Skip to main content
FromNine
Menu

AI & data

Generative AI that holds up in production.

We build generative AI applications on your own knowledge and data, with retrieval that respects permissions, evaluation that gates each release, and cost, quality and latency you can see.

For organisations that run one AI platform across several countries, languages and sets of national rules.

AI Engineering & GenAI: capability stackFour layers stacked in depth. From top to bottom: retrieval over your sources, the model gateway, evaluation and guardrails, and the application your users work in. One route runs from the sources to the application.01Retrieval over your sources02Model gateway and routing03Evaluation and guardrails04Application in production
  1. 01Answers that cite their sources and decline when the evidence is missing
  2. 02A golden dataset and regression gates agreed with your teams before go-live
  3. 03Model routing and budgets that keep cost and latency per request visible

01 Problems

Where generative AI stalls

The model is rarely the hard part. What stalls most programmes is everything around it.

  1. 01

    The pilot impressed, then nobody could say if it was good

    Quality was judged by a few people trying a few questions. There is no test set, so every prompt change or model upgrade is a gamble.

  2. 02

    The assistant shows people documents they should not see

    Content was indexed without its access rights. Security and the works council stop the rollout, and they are right to.

  3. 03

    Costs grow faster than usage

    Every request goes to the largest model with the longest context. Nobody knows the cost of one answer, so nobody can manage it.

  4. 04

    The answers sound right and are not

    Retrieval returns near misses, the model fills the gaps fluently, and users lose trust after the third confident mistake.

02 Delivery ledger

What we deliver

  1. 01 Retrieval and enterprise knowledge systems

    Retrieval designed per source: how content is parsed and chunked, how lexical and vector search combine, and how access rights travel with every passage. The basis for enterprise knowledge assistants.

    What we do

    • Source inventory with owners, formats and access models
    • Chunking and metadata strategy per document type
    • Hybrid lexical and vector search with re-ranking
    • Permission-aware retrieval that inherits source access rights

    What you receive

    • Ingestion pipelines with change detection
    • Retrieval service with citations and abstention rules
    • Retrieval quality report per source
  2. 02 Model strategy without lock-in

    A model gateway between your applications and the models, so you can route by task, fall back when a provider fails and swap models as prices and capabilities change.

    What we do

    • Options across hosted frontier models, EU-hosted deployments and self-hosted open-weight models
    • Routing by task, sensitivity and cost
    • Fallbacks, timeouts, caching and rate limits
    • Data processing terms reviewed with your legal and security teams

    What you receive

    • Model gateway with routing policy as configuration
    • Model decision record per use case
    • Cost and latency budget per request type
  3. 03 Evaluation as an engineering discipline

    Quality criteria agreed with the people who own the process, turned into test sets that run on every change. Model-graded checks are calibrated against human review before we rely on them.

    What we do

    • Golden datasets per use case, built with domain experts
    • Offline evaluation of retrieval and generation separately
    • Red-teaming for prompt injection and data leakage
    • Online evaluation from production feedback and sampled review

    What you receive

    • Evaluation suite wired into CI as a release gate
    • Quality criteria signed off by the business owner
    • Red-team findings and mitigations
  4. 04 LLMOps and observability

    Tracing from question to retrieved passages to model call to answer, with prompts and models versioned like code and budgets enforced in production.

    What we do

    • Tracing and structured logging of every request
    • Prompt, model and index versioning
    • Semantic and exact caching where it is safe
    • Drift and quality monitoring with alerting

    What you receive

    • Dashboards for quality, cost and latency
    • Incident playbook for AI-specific failures
    • Runbook for model and index updates
  5. 05 Governance by design

    We design with the EU AI Act in mind, from risk classification and your role as provider or deployer to logging, human oversight and documentation. Legal assessment stays with your own counsel.

    What we do

    • Use-case classification and obligations by role
    • Input to data protection impact assessments under GDPR
    • Logging and human-oversight design
    • Model and system documentation kept with the code

    What you receive

    • System documentation pack
    • Logging and retention design
    • Oversight procedure for each use case
  6. 06 Fine-tuning and document understanding, when evaluation says so

    Fine-tuning, distillation and multimodal document understanding are tools for specific gaps. We use them when the evaluation suite shows that retrieval and prompting are not enough, and not before.

    What we do

    • Gap analysis from evaluation results
    • Data preparation and labelling plan
    • Distillation to smaller models for cost or latency
    • Layout-aware extraction for scanned and complex documents

    What you receive

    • Comparative evaluation before and after
    • Training data lineage
    • Decision to keep, roll back or retire

03 Architecture

An illustrative reference architecture

Most enterprise generative AI systems we build share one shape: retrieval over your sources, a model gateway, guardrails on both sides of the model, and an evaluation and observability loop underneath that decides what may go live.

The permission filter sits inside retrieval, not in the user interface. If a person cannot open a document in the source system, its passages never reach the model on their behalf.

Layers in the drawing

Sources
Document stores, wikis and business systems, plus the identity provider that knows who may see what.
Retrieval
Ingestion, a hybrid index, re-ranking and the permission filter.
Models
A gateway that routes between hosted, EU-hosted and self-hosted models, with fallbacks and caching.
Application
Cited answers, abstention when evidence is missing, and guardrails on input and output.
Run and improve
Golden sets, regression gates, traces and budgets. Production traces feed the next test set.
Illustrative example
  1. Sources

    • Document stores and wikis
    • Business systems
    • Identity provider
  2. Retrieval

    • Ingestion and chunking
    • Hybrid index
    • Permission filter
    • Retrieval and re-ranking
  3. Models

    • Model gateway
    • Hosted, EU-hosted and self-hosted models
    • Input and output guardrails
  4. Application

    • Application with citations
  5. Run and improve

    • Golden sets and regression gates
    • Traces, cost and latency budgets
Retrieval, model gateway and evaluation loop

Illustrative architecture, not a client system.

Read the diagram as text

The main route runs from document stores through ingestion, a hybrid index, retrieval with re-ranking and the model gateway to the application.

Business systems also feed ingestion. The identity provider feeds a permission filter that constrains retrieval. Guardrails check input and output of the application. The gateway routes to hosted frontier models, EU-hosted deployments and self-hosted open-weight models.

Underneath, an evaluation suite with golden sets and regression gates, and observability with traces, cost and latency budgets. Production traces feed the evaluation suite.

04 Considerations and limits

Engineering considerations and limits

  • Quality is measured, not promised

    How we handle it

    We agree quality criteria with your process owners and test AI output against them before and after each release.

    Retrieval and generation are evaluated separately, so a failure points to its cause.

    Limits and dependencies

    AI output can be wrong. That is why we design review steps for uncertain cases and abstention when evidence is missing.

    A test set is only as good as the experts who help build it. We need their time early.

  • Permissions travel with the content

    How we handle it

    Access rights are captured at ingestion and enforced at query time against your identity provider.

    Limits and dependencies

    Where source systems hold rights in ways that cannot be read reliably, we keep that content out of scope until it can be.

    Over-shared content in the source stays over-shared. Retrieval makes it easier to find, so a clean-up is sometimes the first step.

  • Cost and latency are design inputs

    How we handle it

    Each request type gets a budget. Routing, caching and context size are tuned against it, and the dashboards show where money goes.

    Limits and dependencies

    Provider prices and rate limits change outside your control. The gateway lets you react, but it cannot make a price change disappear.

  • Models change underneath you

    How we handle it

    Models are pinned by version, and every upgrade runs the full evaluation suite before it reaches users.

    Limits and dependencies

    Hosted models are sometimes retired at short notice. We plan an alternative per use case, and that plan needs owner approval.

05 How we work

How we work

The order matters: we decide how to measure before we decide what to build.

  1. 01

    Frame the use case

    Who uses it, which decisions it supports, which sources it may touch and what a wrong answer costs.

    OutputUse-case brief and risk classification

  2. 02

    Build the golden set

    Real questions and expected answers, collected with your domain experts, including the ones the system should refuse.

    OutputVersioned test set

  3. 03

    Engineer retrieval first

    Ingestion, index and permission filter, tuned until retrieval quality meets the agreed bar.

    OutputRetrieval quality report

  4. 04

    Release behind gates

    Generation, guardrails and the gateway go live only when the evaluation suite passes in CI.

    OutputRelease with evaluation evidence

  5. 05

    Operate and improve

    Traces, sampled review and user feedback feed the next test set and the next release.

    OutputMonthly quality, cost and latency review

06 Human control

Where people stay in control

Generative AI drafts, retrieves and suggests. People decide what it may see, what counts as good and what happens when it is unsure.

  • Owners sign off quality

    The business owner of each use case approves the quality criteria and the release evidence, not the delivery team alone.

  • Uncertain answers go to a person

    Below an agreed confidence threshold the system abstains or routes the question to a named team.

  • Access stays with the source

    Who may see a document is decided in the source system and your identity provider, never inside the AI layer.

  • Feedback changes the system

    Users can flag answers. Flags are reviewed by people, and confirmed failures become test cases.

07 Technologies

Technologies we work with

Retrieval
  • OpenSearch and Elasticsearch
  • PostgreSQL with pgvector
  • Azure AI Search
  • Cross-encoder re-rankers
Models
  • Hosted frontier models
  • EU-region managed model services
  • Open-weight models on vLLM
  • Embedding models per language
Evaluation and ops
  • OpenTelemetry tracing
  • CI-based evaluation suites
  • Prompt and model registries
  • Langfuse or equivalent
Application
  • TypeScript and Python services
  • Streaming APIs
  • Existing portals, Salesforce, SAP and Odoo front ends

Listing a technology describes our engineering experience. It does not imply a partnership with or endorsement by its vendor.

08 Example

An illustrative example

Illustrative example

A policy assistant for a group operating in five Member States

01Situation
Group policies exist in English, while local procedures are written in Dutch, French, German and Italian. Staff ask the same questions in five languages and get different answers.
02What we would build
Multilingual retrieval over group and local documents, with access rights inherited from each country’s document store and answers that cite the governing text.
03Where people decide
Each country’s policy owner approves the test set for their language and signs off releases that change local answers.
04What we would measure
Share of answers that cite the correct governing document, abstention rate on out-of-scope questions, and cost per answer per language.

09 Sector lens

In your sector

  • Assistants for caseworkers and citizens that cite the legal source, work in every official language and log what was retrieved for whom.

  • Policy and procedure assistants with strict entitlements, full traceability and model risk documentation your second line can review.

  • Retrieval over protocols and product documentation with clear limits: no clinical decisions, sensitive data kept in approved environments.

International

Many languages, one evaluation standard

International organisations rarely deploy generative AI in one country only. The EU AI Act sets one framework, but data protection authorities, works councils and sector supervisors still differ by Member State, and so do the languages your knowledge is written in.

We design for that in discovery: retrieval and evaluation per language, data residency set per use case, and logging and documentation that serve each country’s questions without a separate system for each. Programmes funded under Digital Europe often ask for exactly this kind of reusable, documented design.

Questions about AI engineering

Which model do you recommend?

It depends on the task, the data involved and your hosting requirements. We usually combine models: a capable model for reasoning-heavy steps, smaller or self-hosted models for classification and extraction, and EU-hosted deployments where data must stay in the EU.

The gateway keeps that choice reversible, so the answer can change as models and prices do.

How accurate will it be?

We do not quote accuracy before we have measured it on your data. We agree quality criteria with you, build a test set, and report results against it before go-live and after each change.

Models make mistakes, so the design always includes what happens when they do: abstention, review steps and a route to a person.

Will our data be used to train someone else’s model?

Not in the architectures we design. We select model services whose terms exclude training on your inputs, or host models ourselves in your environment, and we document the choice per use case with your security team.

Does the EU AI Act apply to our use case?

We help you understand which obligations may apply, based on the use case and your role as provider or deployer, and we design logging, oversight and documentation to match, so your counsel has the facts to reach a conclusion.

How do you handle security and sensitive data?

FromNine is ISO/IEC 27001 certified. Access boundaries, data residency and retention are agreed before we build, and red-teaming for prompt injection and data leakage is part of the release process we set up.

What is a sensible first step?

One use case with a clear owner, a limited set of sources and a test set built with your experts. That gives you a working system and evidence about quality, cost and latency, which is what the next investment decision needs.

Bring one use case and its hardest questions

We will show how we would measure it before we talk about building it.