LLM · agents · MCP

LLM evaluation,
with a record an auditor can read.

Run your own model or agent through a test set and score each answer on nine metrics, judged by a second model. Every run is stored with its inputs, scores and date, which is what ISO 42001 and the EU AI Act mean when they ask for testing that leaves evidence.

01

What it does

  • Nine metrics: relevancy, faithfulness, hallucination, correctness, bias, toxicity, task completion, tool correctness and MCP use
  • Works with Anthropic, OpenAI and OpenAI-compatible endpoints
  • An LLM judge scores each answer and explains the score
  • Each run stored as evidence against AI governance controls
01

What is LLM evaluation?

LLM evaluation is testing a language-model application against a fixed set of prompts and scoring the answers on the properties that matter to you: whether they are relevant, grounded in the source, correct, free of bias and toxicity, and for agents whether the right tools were called with the right arguments. Run on every model or prompt change, it is the regression test for behaviour that unit tests cannot see.

For governance the record matters as much as the score. ISO/IEC 42001 and the EU AI Act both expect testing to be planned, repeated and documented; a dated run with its inputs and results is that documentation.

01

How does LLM evaluation work?

  1. Connect your modelPoint the platform at your model through the Anthropic API, the OpenAI API or any OpenAI-compatible endpoint, including self-hosted models that expose one.
  2. Write test casesEach test case has an input, and optionally an expected answer, the retrieval context the answer should rest on, and the tools or MCP servers the agent may use.
  3. Choose metrics and a thresholdPick the metrics that matter for the case. Each is scored from 0 to 1, higher is better, and the default pass mark is 0.7, which you can change per test case.
  4. Run and judgeThe platform calls your model with the input, then a judge model scores the answer on each metric and gives a one-sentence reason for the score.
  5. Keep the recordEvery run is stored with its input, output, scores, reasons and date, so you can compare runs over time and show the history to an auditor.
01

Which metric answers which question?

Nine metrics, each asking one question. Some need extra inputs, such as an expected answer or retrieval context.

MetricThe question it asksNeeds
Answer relevancyIs the answer on topic for the input?Input and output
FaithfulnessIs every claim supported by the retrieval context?Output and context
Hallucination-freeDoes the answer avoid facts that contradict the context?Output and context
CorrectnessDoes the answer match the expected answer?Output and expected answer
Bias-freeIs the answer free of political, gender, racial or ideological bias?Output
Non-toxicIs the answer free of toxic, harmful or hateful content?Output
Task completionDid the agent do what the input asked?Input and output
Tool correctnessWere the right tools called with sensible arguments?Input and tool calls
MCP useDid the agent use the available MCP tools, resources and prompts correctly?Input, output and MCP primitives

For hallucination, bias and toxicity the score measures the absence of the harm, so higher is better for every metric.

01

Why test MCP tool use separately?

Because an agent can give a fluent answer while calling the wrong tool, or the right tool with the wrong arguments. Tool correctness and MCP use are scored on their own, so a regression in how the agent acts shows up even when its words still read well.

01

Why use a model to judge another model?

A judge model is used because the properties that matter, such as relevance, grounding and tone, cannot be checked with an exact string match. A second model reads the input, the answer and any context, and scores it against a written rubric for one metric at a time.

A judge can be wrong, so every score comes with its reason. A person can read why an answer scored 0.4 on faithfulness and decide whether the judge was right. Treat the scores as a regression signal across runs rather than a single verdict on the model.

01

What does ISO 42001 or the EU AI Act expect from testing?

Both expect testing to be planned, repeated and recorded. ISO/IEC 42001 asks an organisation to operate its AI systems under controlled processes and to verify and validate them, with documented evidence of the results. For high-risk systems, the EU AI Act requires a risk management system (Article 9), appropriate levels of accuracy and robustness (Article 15) and a quality management system that covers testing and validation (Article 17). The NIST AI RMF calls this work the Measure function.

None of these frameworks names a metric or a pass mark. What an assessor looks for is a defined test set, criteria agreed in advance, results kept over time, and action when results fall. A dated run with its inputs, scores and reasons is that record. See ISO 42001, the EU AI Act and NIST AI RMF.

01

When should evaluations run?

Run evaluations whenever something that shapes the answer changes: a new model version, a changed system prompt, a new retrieval index or a new tool. Those are the changes that quietly alter behaviour without breaking any unit test.

Keep the test set stable between runs so scores are comparable, and add a case every time a real failure is found in production. Over time the test set becomes a record of every way the system has gone wrong, and a guard against it happening again.

01

How do you test an agent that uses tools?

Test an agent's actions separately from its words. An agent can write a fluent summary after calling the wrong tool, or the right tool with the wrong arguments. Tool correctness scores the calls themselves, and MCP use scores whether the agent used the tools, resources and prompts its MCP servers offered.

Record the tool calls with each run. When a score drops, the calls show what the agent did differently, which is usually faster to diagnose than the text. The wider governance work, such as inventory, risk classification and oversight, is covered in AI governance.

01

What will an auditor ask for?

  • The inventory of AI systems and which ones are tested
  • The test plan: what is tested, how often and against what criteria
  • Results over time, not a single run before the audit
  • What happened when results fell below the agreed mark
  • Who reviewed the results and approved the release
  • How test cases are added when real failures are found
01

What it does not do

  • It does not decide your pass marks; you set the threshold per test case
  • It does not certify a model or make an AI system compliant on its own
  • It does not run without a judge model configured on the platform
  • It does not generate your test set; test cases come from you
  • It does not replace human review of high-risk outputs
01

Where this sits

Every result here is a control result on the same graph as the rest of the platform, so it reaches each framework that asks for it without being gathered again. Coverage lists the regimes.

Related reading: AI governance, the ISO 42001 guide and the EU AI Act guide.

Questions

The things people ask us

Which models can be tested?

Any model behind an Anthropic, OpenAI or OpenAI-compatible API, including self-hosted models that expose an OpenAI-compatible endpoint.

Who judges the answers?

A second model scores each answer against the metric's definition and records its reasoning, so a score can be reviewed by a person.

Does this make us EU AI Act compliant?

No single tool does. Evaluation evidence supports the testing and accuracy obligations; risk classification, documentation and human oversight are separate pieces of work covered in AI governance.

What is the default pass mark?

0.7 on a 0 to 1 scale, set per test case. Choose a mark you can justify and keep it stable so runs stay comparable.

Do faithfulness and hallucination need retrieval context?

Yes. Both check the answer against the context it should rest on, so supply the retrieved passages with the test case.

Can we test a self-hosted model?

Yes, if it exposes an OpenAI-compatible API. Many self-hosted serving stacks do.

Why does every score have a reason?

So a person can check the judge. A score without a reason cannot be reviewed, and a reviewed score is what an assessor wants to see.

How many test cases do we need?

Enough to cover the main tasks and every known failure. Start small, and add a case each time something goes wrong in production.

Book a walkthrough

Your next audit could be a link.

Thirty minutes. We connect one cloud account live and show you real evidence landing in the ledger before the call ends.