Skip to main content
FromNine
Menu

AI & data

Data that AI and decisions can rely on.

We build data platforms where every dataset has an owner, a contract and a known quality, so analytics, operations and AI all draw on the same trusted numbers.

For organisations that need one trusted view of data spread across countries, entities and legacy systems.

Data & Analytics: capability stackFour layers stacked in depth. From top to bottom: analytics and AI that consume data, data products with owners and contracts, the lakehouse layers that cleanse and conform, and governance across all of it.01Analytics and AI use02Data products and contracts03Lakehouse layers04Governance and lineage
  1. 01Data products with named owners, contracts and quality checks
  2. 02Governance that encodes purpose limitation and retention, not just a catalogue
  3. 03One set of KPI definitions that business and IT both sign

01 Problems

Familiar data problems

Most AI projects that disappoint turn out to be data projects in disguise.

  1. 01

    Three reports, three different numbers

    Finance, operations and the board each calculate revenue or caseload their own way. Meetings start with reconciling instead of deciding.

  2. 02

    Nobody owns the data, so nobody fixes it

    Errors are found downstream, corrected in a spreadsheet and come back next month.

  3. 03

    The AI team spends most of its time finding data

    Every use case starts with weeks of access requests, extracts and cleaning that the next team will repeat.

  4. 04

    Archives are unusable for AI

    Scanned documents with poor OCR, duplicates and no metadata. Retrieval over them returns noise.

02 Delivery ledger

What we deliver

  1. 01 Data foundations for AI and analytics

    Lakehouse architectures with layered modelling, from raw to conformed to consumption, and data products that publish what they contain, how fresh it is and who answers for it.

    What we do

    • Platform architecture and storage format choices
    • Layered modelling conventions
    • Data product design with contracts
    • Batch and change data capture ingestion

    What you receive

    • Lakehouse platform defined in code
    • First data products in production
    • Data contract template and review process
  2. 02 Quality, lineage and master data

    Automated quality checks at every layer, lineage from source to report, and master and reference data for the entities that matter: citizens, customers, products and assets.

    What we do

    • Quality rules with thresholds and alerts
    • Column-level lineage capture
    • Master data matching and survivorship rules
    • Reference data management

    What you receive

    • Quality dashboard per data product
    • Lineage you can query
    • Golden records for priority entities
  3. 03 Governance that is enforced

    A catalogue, classification and access policies that the platform applies automatically. GDPR purpose limitation and retention are encoded, and data sharing follows the rules of the Data Governance Act where it applies.

    What we do

    • Classification scheme and tagging
    • Attribute-based access policies
    • Retention and deletion automation
    • Data sharing agreements and technical controls

    What you receive

    • Catalogue populated from the platform
    • Access policies as code
    • Retention schedule in operation
  4. 04 Unstructured content made ready for AI

    OCR quality improvement, metadata enrichment, deduplication and archive clean-up, so document intelligence and retrieval systems start from content worth retrieving.

    What we do

    • Content inventory and quality sampling
    • OCR and layout extraction improvements
    • Near-duplicate detection
    • Metadata enrichment and classification

    What you receive

    • Cleaned, enriched content store
    • Content quality report
    • Rules for content added from now on
  5. 05 Analytics and BI people use

    A semantic layer with KPI definitions agreed by business and IT, governed self-service and embedded analytics inside SaaS products and portals.

    What we do

    • KPI definition workshops
    • Semantic model and metrics layer
    • Self-service guardrails and certified datasets
    • Embedded analytics for customer-facing products

    What you receive

    • KPI glossary with owners
    • Certified semantic model
    • Dashboards that replace spreadsheets
  6. 06 Streaming for operational use

    Change data capture and event streams where decisions cannot wait for the nightly load: stock levels, case status, fraud signals.

    What we do

    • Latency requirements per use case
    • CDC from operational databases
    • Stream processing and materialised views
    • Replay and backfill design

    What you receive

    • Streaming pipelines in production
    • Operational views with freshness targets
    • Replay procedure

03 Architecture

An illustrative reference architecture

Data moves from operational systems and documents through ingestion into a lakehouse with layers: raw as received, then conformed and quality-checked, then data products with an owner and a contract. Consumers never read raw data.

From the data products, one route feeds the semantic layer and BI; others feed AI retrieval and features, and data APIs for other systems. Governance sits under all of it and is applied by the platform, not by a policy document.

Illustrative example
  1. Sources and ingestion

    • Operational systems
    • Documents after OCR and enrichment
    • Batch and change data capture
  2. Lakehouse

    • Raw layer
    • Conformed layer
    • Quality checks and lineage
    • Master and reference data
    • Data products with contracts
  3. Consumption

    • Semantic layer and BI
    • AI retrieval and features
    • Data APIs and sharing
  4. Governance

    • Governance: catalogue, classification, access policies, retention
Lakehouse layers into owned data products

Illustrative architecture, not a client system.

Read the diagram as text

The main route runs from operational systems through ingestion with batch and change data capture, a raw layer, a conformed layer and data products, to a semantic layer with governed BI.

Unstructured documents also enter through ingestion after OCR and enrichment. Quality checks and lineage cover the raw and conformed layers. Master and reference data feed the data products. Data products also serve AI retrieval and features, and data APIs for sharing.

A governance band with catalogue, classification, access policies and retention spans the platform.

04 Considerations and limits

Engineering considerations and limits

  • Ownership before tooling

    How we handle it

    Each data product gets a business owner who decides what it means and a technical owner who keeps it running. We help define both roles and the contract between producers and consumers.

    Limits and dependencies

    A platform cannot create ownership. If the organisation will not assign owners, quality problems will return, whatever the tooling.

  • Purpose limitation in practice

    How we handle it

    Personal data is tagged with the purposes it may serve, and access policies and retention follow those tags automatically.

    Limits and dependencies

    Which purposes are lawful is a decision for your data protection officer and counsel. We implement it; we do not decide it.

  • Data for AI has its own quality bar

    How we handle it

    For AI use cases we measure what retrieval and models need: coverage, freshness, duplicates and the quality of text extraction.

    Limits and dependencies

    Some archives are not worth fixing. A sample-based assessment will tell you, and sometimes the right answer is to leave them out.

05 How we work

How we work

We build the platform around the first data products that matter, not the other way round.

  1. 01

    Pick decisions, not datasets

    Start from two or three decisions or AI use cases and the data they need, with owners named.

    OutputData product backlog with owners

  2. 02

    Build foundations around them

    Platform, ingestion, quality checks and governance, just enough to serve those products properly.

    OutputFirst data products in production

  3. 03

    Grow by product

    Each new product reuses the foundation; the catalogue and the KPI glossary grow with it.

    OutputProduct roadmap and adoption measures

06 Human control

Where people stay in control

Data platforms automate movement and checks. People decide meaning, access and purpose.

  • Owners approve definitions

    A KPI or data product changes meaning only when its business owner approves, and the change is versioned.

  • Access is granted by policy owners

    Data owners decide who may use which data for which purpose; the platform enforces it and logs every grant.

  • Quality breaches stop the line

    When a contract check fails, consumers are told and the owner decides whether to publish, hold or correct.

07 Technologies

Technologies we work with

Platforms
  • Databricks
  • Microsoft Fabric
  • Snowflake
  • Open lakehouse on Apache Iceberg or Delta Lake
Pipelines
  • dbt
  • Apache Spark
  • Kafka and Debezium
  • Cloud-native orchestration
Governance
  • Unity Catalog and Purview
  • Open-source catalogues
  • Data quality frameworks
  • Policy-based access control
Analytics
  • Power BI
  • Semantic and metrics layers
  • Embedded analytics
  • SAP analytics (existing estates)

Listing a technology describes our engineering experience. It does not imply a partnership with or endorsement by its vendor.

08 Example

An illustrative example

Illustrative example

One set of operational KPIs for a group in several countries

01Situation
Each country reports throughput and backlog from its own system with its own definitions. Group management compares numbers that do not mean the same thing.
02What we would build
Conformed data products per country feeding one semantic layer with KPI definitions agreed at group level and local detail preserved.
03Where people decide
Country controllers approve the mapping of their data; the group KPI owner approves definitions and changes.
04What we would measure
Time to produce the monthly report, number of reconciliation corrections and use of certified dashboards.

09 Sector lens

In your sector

  • Governed research and operational data with strict purpose limitation, pseudonymisation and access logging.

  • Product and asset master data, quality and production data joined up for planning, traceability and AI.

  • Authentic sources, once-only data use and reporting that policy makers and auditors can trace back to the record.

International

Data that crosses borders

Groups and public bodies that work in several Member States share data under rules that look alike but are not the same. The GDPR sets one framework for personal data, the Data Act gives users rights over data from connected products and services, and national interpretations of retention and purpose still differ.

We design data products whose contracts include where data may be processed, for which purposes and for how long, so the platform can apply the right rule per country without separate copies of the data.

Questions about data and analytics

Do we need a data mesh?

You need the useful parts of it: data products with owners and contracts. Full decentralisation pays off in large organisations with mature domain teams. Many organisations do better with a central platform team and domain ownership of the products.

Is our data ready for AI?

A short, sample-based assessment of the sources a use case needs will tell you: coverage, quality, access rights and legal basis. The answer is usually “partly”, with a clear list of what to fix first. Our AI engineering teams use the same assessment.

We already have a data warehouse. Do we start again?

Rarely. A warehouse that serves trusted reports is an asset. We usually extend it with lakehouse capabilities for unstructured and high-volume data, and migrate only where cost or limits justify it.

How do you handle data sharing with partners or other authorities?

Through data products exposed via governed APIs, with agreements, purpose limitation and logging in place. Where the Data Governance Act or the Data Act apply, we design the technical controls to match.

Name the decision you cannot make with today’s data

We will trace what it needs back to the sources and show what a first data product would take.