Service 07

Evals, security and AI governance

We measure whether your AI systems work, test how they fail and document them for auditors, with evals, red-team tests, tracing and compliance mapping.

In one paragraph

An eval is a repeatable test that measures how well an AI system does a specific task. An eval suite runs a set of real cases through the system and scores each result against a written standard.

Deliverables

What you get

  1. 01 Offline eval suites built from real cases, with clear pass criteria
  2. 02 Online monitoring that samples production runs and alerts on quality drops
  3. 03 Tracing based on the OpenTelemetry semantic conventions for generative AI
  4. 04 Red-team tests for prompt injection, data leaks and tool misuse, based on the OWASP Top 10 for Agentic Applications
  5. 05 An inventory of your AI systems with an EU AI Act risk classification
  6. 06 An ISO/IEC 42001 gap assessment and a map to the NIST AI RMF

Where it fits

Typical use cases

  • Before launch

    An eval suite and a red-team report that show the system is ready, or exactly what blocks it.

  • Model upgrades

    A regression run on every model or prompt change, so a new version cannot quietly make results worse.

  • Audit and compliance

    Documentation, risk classification and controls that your auditors and regulators can follow.

  • Incident review

    A trace of exactly what the agent saw, decided and did, and a new test that stops the failure from coming back.

How we work

Evals come first

We write the evals before we build the system. An eval set starts with 50 to 200 real cases and a written definition of a correct result for each. It grows every time production shows us a new failure.

Some checks are exact, such as the right account or the right amount. Some need a grader: a rubric scored by a model, then checked against human scores until they agree.

Test how it fails

Agents read untrusted input, such as emails, web pages and documents, and they can act. That combination needs its own security testing. We test for prompt injection, data leaks across users, tool misuse and runaway cost. Then we fix the design, not just the prompt.

Know what runs where

Governance starts with an inventory. You need to know which AI systems you run, who owns each one, what data it uses, what it can do and which rules apply to it. We build that inventory, classify each system and set up the controls and the documentation to keep it current.

Frameworks we map to

  • The EU AI Act and its risk classes.
  • ISO/IEC 42001, the standard for an AI management system.
  • The NIST AI Risk Management Framework and its generative AI profile.
  • The OWASP Top 10 for Agentic Applications.

FAQ

Questions about evals and governance

We already have a working pilot. Why add evals now?

Without evals you cannot tell whether a change made the system better or worse. That includes a new prompt, a new model or a vendor update. Evals turn each change from a guess into a measurement.

Can prompt injection be fully prevented?

No. Nobody can fully prevent it today. We design for containment instead. Untrusted content is isolated, tools have narrow scopes, irreversible actions need approval and outputs are checked before they leave the system.

Does the EU AI Act apply to us?

It can apply if you place AI systems on the EU market or if their output is used in the EU, even when the company is outside the EU. We map each system to the risk classes and the obligations that apply to it, and we work with your legal counsel on the conclusions.

Related

All services
  • Service 01

    AI agents in production

    We design, build and run AI agents that work inside your systems. They read, decide, call tools and hand the case to a person when a rule says so.

  • Service 08

    AI strategy and operating model

    We help leadership pick the AI work that pays off, decide what to build or buy, and set up the operating model that keeps AI systems owned and measured.

Tell us which process you want to hand to an agent

A 30-minute call with an engineer. We will tell you whether an AI system is the right tool for it, and what it would take to run it in production.