October 1, 2026

LLM-as-a-Judge for GenAI Systems

A practical evaluation framework for delivery teams on Databricks

1. An explanation of what LLM-as-a-Judge is and for Whom it is intended.

The LLM-as-a-Judge framework is the one that we apply in order to assess generative AI systems prior to them being released and also to continuously check them after they have gone live; it covers RAG pipelines, customer support bots, knowledge assistants, text-to-SQL tools, and the different copilots that we build for our internal teams. If you have one of these on Databricks, the document here explains how you can determine whether it is actually working.

We have a dedicated LLM that assesses the outputs of our generator LLM against clearly defined criteria, records the scores in MLflow along with the rationale, and uses these scores to decide when to release new versions and to monitor production. The remainder of the document covers the metrics, how they have been implemented, the thresholds we at V4C suggest starting with, and the position this approach holds within a typical Databricks GenAI stack.

2. Why do we need an LLM judge

We need an LLM judge because large language models can review lots of information quickly and fairly. They do not get tired or let their feelings affect their decisions. This helps make the judging process more steady and less biased. LLM judges can also check work in the same way every time, making results more reliable. In fields with large volumes of data or many entries, using an LLM judge saves time and effort.

When assessing a GenAI system as one would traditional software, three things fail, and these failures add up.

The output is not deterministic; for example, the question ‘What’s our refund policy?’ could be answered in three different ways, all of which would be correct. Using BLEU, ROUGE, or F1 scores based on surface similarity to a reference fails when there is no single standard reference, and the situation is made worse by the fact that a fluent hallucination can score highly on token overlap simply because it employs the appropriate vocabulary. As a result, you get misleading metrics and faulty systems.

The scale of volume goes beyond what humans can cope with. A relatively small production assistant deals with tens of thousands of queries each month, while the bigger ones handle even more. It’s not possible, within any reasonable budget, to have a domain expert read every response or even a substantial number of them. If you don’t have an automated evaluator, your only alternatives are to randomly check some responses and hope for the best, or to release the product without any measurement and then find out from users.

There is no measurable baseline for any change. If a team alters a prompt, changes the embedding model, or increases the top-k value, the demonstrations still seem correct. However, three weeks later, a customer reports a regression that has been slowly degrading retrieval since the change. We’ve all experienced this situation. The solution isn’t to use more careful prompting; it’s to have an evaluator attached to every change so that regressions appear before they do for users.

The LLM judge considers all three aspects: it reviews the response, checks the source material, and then assesses whether the answer is valid before giving a numerical score along with a written explanation. The score is the one that appears on the dashboard or triggers an alert. The explanation is what makes the system auditable: when a metric decreases, the explanation tells you why; and if a regulator or an internal reviewer asks how it was determined that the system was working, the explanation can be cited in the review. Research carried out using MT-Bench and similar benchmarks has consistently found that well-prompted LLM judges agree with expert human raters to a degree comparable to the way humans agree with one another, so they are good enough to replace manual review on a large scale, as long as you periodically calibrate the system against human judgments.

It is important to make clear what advantage this provides: it doesn’t just give you a good story; it also gives you speed. When a judge is connected, all changes are monitored, and the team is able to release updates more quickly since regressions are detected at the time of the PR rather than having to wait until they appear in the support queue. Without such a system, the team ships more slowly and has to learn from customers.

3. The role of the judge within a Databricks GenAI architecture

The same structure is generally followed by the various GenAI systems on Databricks. This is where the judge inserts themselves:

User query
(chat, API, app)
→ Retriever
Vector Search
→ Generator
Mosaic AI Agent
→ Response
to user
→ LLM Judge
MLflow Eval

‍

The judge is applied as a step coming after the traced outputs have been produced. There are two modes:

  • Offline: Before you promote a new version of the model, you have to run the judge against a curated evaluation set. That is your release gate.
  • Online: MLflow Tracing records production traffic, and a sample is scored on a schedule using the same scorers and code as when running the logs from yesterday.

Since the same scorer definitions are used in both cases, the offline thresholds you set are directly carried over into your production, and there is no need for you to calibrate two different evaluators.

Databricks features that pair naturally with this:

  • Mosaic AI Agent Framework for the generator side
  • MLflow Tracing for capturing every production call (and its retrieval context)
  • MLflow Evaluation for running judges over those traces
  • Unity Catalog for storing eval datasets and judge configs as versioned artifacts
  • Lakehouse Monitoring for alerting when scores drift below your thresholds

Think of the eval dataset as a first-class asset: version it, store it in Unity Catalog and check for changes the same way you would examine changes to production code.

3.1 Exactly what MLflow Tracing captures (and the reason span types are important)

Tracing isn’t a single, vague entity; each step of an agent run is recorded as a typed span, and some of the judges require those types to be present. The span types you’ll most frequently encounter:

  • The root span refers to the agent's initial call.
  • RETRIEVER: This refers to the process of carrying out a search within a vector search or any other retrieval system. It is necessary in order to ensure groundedness, retrieval relevance, and retrieval sufficiency, since without a RETRIEVER the judge has no basis on which to assess the entries.
  • Tool: The function or tool calls, including their inputs and the return values.
  • CHAT_MODEL: The language model makes calls, using prompts, responses, and parameters.
  • A chain is a series of operations, something that is typically used when working with LangChain or similar wrappers.

Actual consequence: if the judges on the retrieval side are quietly generating nulls or always-pass scores, then check the trace first. In nine out of ten cases, the retriever is not emitting a RETRIEVER span (a common reason being that it isn’t instrumented or that it’s contained within a general CHAIN span). The judge is not the problem; the trace is.

3.2 The feedback loop between production and evaluation

If you deploy using the Mosaic AI Agent Framework, Databricks will automatically turn on AI Gateway-enhanced inference tables. These are tables within Unity Catalog that are queryable via SQL and record each request, response, latency, model version, and trace from your deployed agent, without any additional instrumentation.

This is important for evaluation because it establishes a connection between offline and online activities. If you run a query on yesterday’s inference table, filter the results to include cases with low-confidence answers, thumbs-down feedback, or instances of specific failure patterns, you will have some new examples that you can add to your evaluation set. The cycle is as follows:

  1. Use the ship with offline evaluation as the gate.
  2. Generate traffic logs and automatically populate the inference tables.
  3. Each week, check the tables for any failures, edge cases, and drift.
  4. Include interesting examples in the offline evaluation set (with the correct answers added).
  5. Before the next release, rerun the offline evaluation, then continue doing this.

Without the loop, your evaluation set will quickly become outdated, since it reflects what users were expected to ask six months ago, not the questions they are actually asking now.

4. The Important Metrics

Eight scorers cover most of what we ship. You don’t need all eight on every project; however, you must provide a specific business justification for any exclusions. 

4.1 The judge taxonomy: select the appropriate tool and then choose the metric

It’s useful beforehand to become familiar with the different types of judges. MLflow offers four levels, the extent of customization and the amount of effort increasing with each level:

Approach Customization When to use
Built-in judges Minimal Scorers for LLMs that have been validated by research, such as Correctness, RetrievalGroundedness, and Safety; these cover most of what most teams need.
Guideline judges Moderate There is a built-in judge which checks the responses against the natural-language rules that you define, and this comes in two forms: global (meaning the same rules are used for the entire evaluation set) and per-row (which means different rules are applied to each example).
Custom LLM judges Full You draw up the prompt, select the model, and then parse the output; regarding domain-specific criteria, the built-in options fail to handle regulated language, field-specific correctness criteria, or logic that accounts for multiple signals.
Code-based scorers Full Python functions that are decorated with @scorer. They are most suitable for deterministic checks such as format validation, length bounds, exact matches, and latency thresholds. There is no need to call an LLM, there is no cost involved, and the results are not unreliable.

‍

With that framing, here are the eight built-in metrics we lean on most:

Correctness

Did the response state the correct thing? The judge checks the response by comparing it with either the expected answer or a list of expected facts. This is your regression-test signal. If a prompt refactor causes the correctness to drop from 0.91 to 0.78 over your eval set, that’s the alert.

In practice, it picked up the fact that an internal HR bot had begun responding to the question “what’s the parental leave policy?” with last year’s policy after a vector index rebuild failed to pick up the updated PDF, and the system’s correctness check had identified this before any employee did.

Groundedness

Does the response adhere to the actual content of the retrieved context? This is what the hallucination detector checks for. Groundedness should decrease significantly if the model mentions a figure or fact that is not present in the source documents.

In practice, a finance assistant had the audacity to create an SLA tier (“standard support includes a 4-hour response”) that was not included in any of the contract sections. The groundedness score assigned that response a value of 0.2, and the rationale noted that the source was missing.

Retrieval Relevance

Does the retriever actually return chunks that are relevant to the question? This allows the retriever problems to be separated from the generator problems, since in practice most of the time spent debugging is devoted to that separation.

In practice, when you reindex with a new embedding model and one of the chunking settings fails silently, retrieval relevance collapses, and all subsequent steps appear worse, though for the wrong reason. Because this metric is missing, you end up spending a whole day adjusting your prompts.

Retrieval Sufficiency

Even if the chunks do happen to be relevant, are they sufficient to answer the question? A retriever might be on topic but still lacking in depth.

In practice, it identifies multi-hop questions in which the first chunk is relevant, but the actual answer requires a second chunk that did not appear in the top-k; in this case, you should increase k or reconsider the chunking method before attributing the problem to the generator.

Answer Relevance

Did the response really answer the question? Generators have a habit of padding things out. A model can produce a beautifully reasonable paragraph which still fails to get to the point.

In practice, the response to “How do I reset my password?” is a 200-word article on best practices for account security that fails to mention the reset link at all. It is well-grounded and accurate, yet of no practical use.

Completeness

When dealing with questions that have several parts, has each part been answered? The usual mistake is for the model to answer the first sub-question and then forget the others.

In practice, the questions people ask include “What is the refund period and how do I file a claim?” The model provides a detailed explanation of the 30-day window but never outlines the claims process. The user then returns five minutes later with another question, increasing the number of support tickets.

Safety

It must comply in all respects with regard to toxicity, bias, prompt injection, and policy violations if it is to be used in a customer-facing context.

In practice, it detects adversarial inputs that cause the model to generate content that violates the content policy; this should be scored on every production trace, not only on the offline evaluation set.

Instruction Adherence

Did the response adhere to the format and the constraints set out in the system prompt? For example, ‘Reply only as JSON’, ‘use bullet points’, and ‘don’t mention competitors’.

In practice, it picks up downstream parsing failures. For example, if your application expects JSON but the model wraps it in markdown code fences 1% of the time, instruction adherence is what flags this before it appears in a customer ticket.

4.2 Guidelines for the judges: when business rules should be expressed in natural language

The system includes metrics for the main quality aspects, but each project has rules which cannot be easily related to correctness or groundedness, for example ‘never mention the names of competitors’, ‘don’t commit the company to a specific timeline’, ‘always cite the relevant policy section’, and ‘match the brand voice’. For such rules, Guideline judges should be created.

Two flavors, both first-class in MLflow:

Global guidelines (Guidelines class)

The same rules apply to each of the examples in the eval set since they are meant to keep a consistent tone, style, and compliance with policy constraints.

This is important in practice, since business stakeholders do not need to know Python to draw up guidelines. If Compliance sends you a Word document with a set of regulations, you can incorporate it directly into the evaluation pipeline. Changes involve no code updates; they are simply changes to the text. Moreover, the judge's reasoning is self-explanatory since the recommendations are in natural language, unlike raw accuracy figures.

Be specific. "The response must cite the source document" is a usable guideline; "the response should be credible" is not — it gives the judge nothing observable to check. Write guidelines against features you can actually see in the response, and test each one on a handful of known-good and known-bad examples before you trust it across the whole dataset.

4.3 Code-based scorers: in the case of a mechanical check

It is not necessary to use an LLM for some of the checks. Checks relating to length, JSON validity, the presence of citation markers, compliance with regex format, and adherence to latency thresholds are all deterministic, immediate, and free to perform. A simple Python function can be turned into a first-class MLflow scorer by using the @scorer decorator, just as the built-in functions are, and then inserted into MLflow.genai.evaluate().

from mlflow.genai.scorers import scorer
from mlflow.entities import Feedback
 
@scorer
def response_length(outputs):
    """Verify response length is within an acceptable range."""
    word_count = len(str(outputs.get("response", "")).split())
    if 20 <= word_count <= 100:
        return Feedback(
            value="yes",
            rationale=f"Response length ({word_count} words) is appropriate",
        )
    return Feedback(
        value="no",
        rationale=f"Response is too {'short' if word_count < 20 else 'long'} ({word_count} words)",
    )

Wherever possible, use code-based scorers. They don’t drift, are less expensive than LLM judges, and provide the same audit trail as LLM-based scores in the reasoning category.

5. One query that covers everything

This is the kind of trace that a release gate would examine. Work through one question entirely.

Question: “If I open a digital download, is it then possible to obtain a refund?”

Context retrieved (top 3 chunks):

Physical items that have not been opened can be refunded within 14 days of purchase.

After the download link has been accessed, digital downloads cannot be refunded.

Customers can ask about refunds by emailing support at support@example.com.

The response was that you can get a refund within 30 days of purchase by contacting support.

Judge scores: 

Metric Score Judge rationale
Correctness 0.10 The response violates company policy as refunds are not available for digital downloads once they have been accessed.
Groundedness 0.05 The figure of 30 days is not present in any of the retrieved chunks; the chunks state a period of 14 days and only in the case of physical items.
Retrieval Relevance 0.85 The retrieved chunks are relevant to the question.
Retrieval Sufficiency 0.90 The information required to answer correctly is included in the context.
Answer Relevance 0.70 The reply addresses the subject but contains incorrect information.
Completeness 0.60 Failing to include the conditional relating to the download link being accessed.
Safety 1.00 No safety issues detected.
Instruction Adherence 1.00 The format and tone are in keeping with the system prompt.

‍

How to read this. Relevance and sufficiency are both strong, and retrieval performed as intended. The information in the pieces was correct. The model created a 30-day timeframe and disregarded the rule regarding accessed downloads, which is the generator’s fault. The prompt or model, not the retriever, has to be fixed. You would waste a day adjusting chunk size without the metric divide.

Now consider the inverse: A query scoring 0.95 in groundedness, 0.40 in retrieval relevance, and 0.30 in accuracy. The user still gets an incorrect response, but because the generator faithfully processed bad context, the fault lies entirely with the retriever. This is exactly why we track these specific metrics: an identical failure rate can require an opposite diagnosis. 

6. Implementing it on Databricks

6.1 Setup

Install MLflow with the Databricks extras. From a terminal:

pip install --upgrade "mlflow[databricks]>=3.1.0"

Or from a notebook:

%pip install --upgrade "mlflow[databricks]>=3.1.0"

dbutils.library.restartPython()

6.2 The creation of the evaluation dataset

It is the dataset that gives your evaluation its real value. For a good evaluation set, it is essential that it contains ground truth where this is important, is diverse, and is representative.

Principles of composition

For the representative: mimic the way real questions are distributed. If policy lookups account for 60% of production traffic and multi-step reasoning for 10%, your evaluation set should not be 50/50.

Examples of intentionally included cases are questions that are unclear, questions for which the answer does not exist in the corpus, questions using adversarial phrasings, and questions in unexpected forms. As these are the situations in which failures occur, it is necessary to include them carefully.

It varies with respect to length, complexity, field of study, and the assumed skill level of the person asking.

Where ground truth comes from

Expert labeling is the most accurate but also the most expensive option; it should be used in the core part of the dataset and in high-stakes areas.

The production of synthetic content: after generating the expected responses using an LLM, have a human carry out a spot-check. This approach is useful in the initial stages, but don't place any trust in it without conducting an evaluation.

Mining is done with a view to production. The real queries are removed from the inference tables, the important ones are marked, and then they are brought back in. The loop described here, which keeps the evaluation set up to date, comes from Section 3.2.

Schema

Your dataframe needs these columns: 

Column Purpose
request The user query as it came in.
response What your generator produced.
expected_response The ground-truth answer, or a list of expected_facts if there’s no single canonical phrasing.
retrieved_context The chunks the retriever returned for this query. Required for groundedness, relevance, and sufficiency.
row_id UUID for joining results back to your source data.

‍

import pandas as pd
import uuid
 
eval_df = pd.DataFrame({
    "request": [...],
    "response": [...],
    "expected_response": [...],
    "retrieved_context": [...],
})
eval_df["row_id"] = [str(uuid.uuid4()) for _ in range(len(eval_df))]

6.3 Defining your judges

from mlflow.genai.scorers import (
    Correctness,
    RetrievalGroundedness,
    RetrievalRelevance,
    Safety,
)
judge_model = "anthropic:/claude-3-5-sonnet-latest"
correctness  = Correctness(model=judge_model)
groundedness = RetrievalGroundedness(model=judge_model)
relevance    = RetrievalRelevance(model=judge_model)
safety       = Safety(model=judge_model)

6.4 Running the evaluation

Use mlflow.genai.evaluate() as it’s the GenAI-specific entry point and natively understands the scorer types above.

import mlflow
 
mlflow.set_experiment("/Users/<your-user>/genai-eval")
 
results = mlflow.genai.evaluate(
    data=eval_df,
    scorers=[correctness, groundedness, relevance, safety],
    # predict_fn=my_agent.predict,   # optional: regenerate outputs in-line
    # model_id="models:/my-app/3",   # optional: pin to a versioned app
)

‍

Regarding the API: Even if you're using older code that invokes mlflow.evaluate(… model_type="databricks-agent"), it will still work, but the current approach is mlflow.genai.evaluate() and this should be the one that new code aims to use. It includes the same scorers, has GenAI-aware orchestration, and provides a cleaner result format. 

6.5 Pulling scores back out

In MLflow 3, per-example results are stored in traces rather than in an attribute called result_df. Use search_traces to retrieve them. 

import mlflow
 
eval_traces = mlflow.search_traces(run_id=results.run_id)
# Each trace carries its assessments (the judge feedback) alongside the spans
# you can use to debug retrieval and tool-call behavior.

‍

All the other items, such as the runs, the aggregated metrics, the judge rationales, and the original trace tree, are automatically logged to MLflow. To obtain the rationales, do exactly as you do when retrieving the scores, selecting the rationale field rather than the value field. Save both the score and the rationale. The score serves as your dashboard while the rationale acts as your audit trail. 

7. Selecting a judge model

Do not use the same family as your generator; if you are using Claude, then assess the output using GPT-4o or Gemini, and if you are using GPT-4o or Gemini, assess it using Claude. Judges who belong to the same family generally award same-family-generated outputs higher scores. Upon our examination, we discovered inflation of between 5 and 10 points.

Set the temperature at 0.0. The scores you are interested in are the repeatable ones. For the purpose of trend tracking, a judgment that varies by ±0.15 across identical runs is ineffective, since it is not possible to tell whether a 4-point drop is due to noise or a real regression.

You may not realize how important cost is. In a typical RAG evaluation set, 500 examples multiplied by 8 metrics and then by two passes (one offline and one production sample) results in 8,000 judge calls per cycle. For most applications, mid-range models, such as Claude Sonnet or GPT-4o-mini, offer the best cost-to-quality ratio. Larger models should only be used in high-stakes situations where the soundness of the logic is essential. 

8. Running it in practice

  • Since only one figure is unreliable, it is necessary to use multiple measurements. Relying on correctness alone won't reveal whether the model omitted half the question, performed poorly on recovery, or produced hallucinated answers. The collection of metrics is meant to allow triangulation.
  • Make sure everything is versioned: The judge model, the prompt, the evaluation dataset, and the scorer configuration. You need to decide if your system or your evaluator has changed when a metric changes.
  • Keep the rationales: We’ve had two compliance reviews where the rationale text was what satisfied the reviewer and not the score.
  • Keep calibrating using human judges. Once per quarter, a domain expert should assess 50 graded outputs. On the other hand, you should consider the judge: if the level of agreement drops below about 80%, either adjust the judge prompt or change the judge model. 

9. Thresholds and Governance

Here’s where we land on starting thresholds. These aren’t arbitrary; they’re calibrated against what we’ve seen across customer-facing deployments.

Metric Recommended floor Why this number
Correctness ≥ 0.85 Calibration runs indicate that users become aware of the errors and lose trust in the system, and when the figure is above 0.85, most teams say that the bot seems reliable.
Groundedness ≥ 0.90 Intentionally prioritized accuracy. An answer that is wrong but based on sound reasoning at least allows for diagnosis; an answer that is not grounded is a hallucination, and that exacerbates trust problems.
Retrieval Relevance ≥ 0.80 If the result falls below a certain threshold, it usually indicates that the embedding model or the chunking configuration should be examined before you continue tuning the generator.
Answer Relevance ≥ 0.85 Below it, downstream UX feedback quickly turns negative, even when the facts are accurate, because users dislike a bot that fails to answer their question.
Safety violations ≤ 1% For customer-facing systems, 1 in 100 is about as high as you can go. With tools that are used only within an organization, a bit more relaxation is possible, but 5 percent is still the maximum that can be considered acceptable anywhere.

‍

They are merely starting points, not actual rules; in fields involving high stakes, such as law, medicine, or regulated financial advice, the requirement for correctness and reliability is set at 0.95 together with an explanation of why this level has been chosen. In the case of internal experimental tools, however, you can be more relaxed. The key is to state the threshold along with the reason for it, rather than leaving it as a placeholder. 

9.1 When a threshold is crossed

  1. Check the retrieval first. If relevance isn't worthwhile, then the retriever has been changed.
  2. Check the prompt to see if the latest release includes any updates.
  3. Take the earlier version of the model and compare it with this one; if there is a sharp decrease in accuracy or safety, then revert.
  4. Immediately revert and conduct an offline investigation if any safety violations exceed the limit. In production, do not carry out debugging. 

9.2 Things to have on hand for audits

If an internal review or a regulator asks how you know that the system was functioning, then the judge settings (including the model and the prompt), the MLflow experiment with its timestamps, the run history, and the version of the evaluation dataset used for each release should all be readily available. 

10. Restrictions

LLM judges are not infallible; they lack expertise in specialized subjects, operate probabilistically, and absorb biases from their training data. A general-purpose judge may make mistakes when assessing outputs in particular edge cases involving tax law, but only in circumstances different from those of your generator. Modify the judges when the subject area requires it, calibrate them against real people, and keep people informed in any case with legal or medical implications.

QA is achieved through this framework; in difficult cases, it doesn't take the place of expert review, but is used not as a replacement for judgment but as a force multiplier for the team. 

11. How v4c.ai Can Help

v4c.ai is a pure-play Databricks services partner with 750+ certifications, 500+ practitioners, and 200+ enterprise clients across Financial Services, Retail, Manufacturing, and Healthcare. One of the production-grade observability and assessment frameworks that our data engineering and AI practice focuses on when building within the Databricks environment is the LLM-as-a-Judge technique. 

‍

With our support, your team will be able to close the feedback loop, standardize metrics and governance, operationalize evaluation, and upskill your engineering teams. We help your company transform GenAI evaluation from a reactive obstacle into a quantifiable, scalable competitive advantage by addressing the gap between framework creation and scalable platform operations. 

12. Important Lessons

  1. Use a wide range of measures, separate the process of creating from that of retrieving, and place the justifications beside the scores.
  2. Make sure that anything which comes into contact with the evaluation pipeline is versioned and use a deterministic, cross-family judge.
  3. Rather than using placeholders, make your thresholds explicit and provide a rationale for them.

Take this step before deploying any generative AI system.

‍

Let’s Get Started
Ready to transform your data journey? v4c.ai is here to help. Connect with us today to learn how we can empower your teams with the tools, technology, and expertise to turn data into results.
Get Started