Blog
UnlikelyAI
Any AI system in production needs a way to check its outputs. One way of doing that is having LLM-as-Judge, which is when you have one language model evaluate the output of another. This method is becoming more common because it is fast, cheap and it allows to scale the system without having a human reviewing the data and labelling the outputs.
The risk is that a judge which works brilliantly in a demo turns inconsistent and inaccurate under real conditions. With the new EU AI Act now in force, you cannot afford answers you won’t be able to defend when a regulator asks how the model reached them. So the question is: how do you keep the benefits while staying reliable in production?
For anyone building these guardrails, Dr Ailish McLaughlin (Solutions Lead) and Alex Corsham, who builds UnlikelyAI’s guardrails, walk through how LLM-as-a-judge works, its benefits and its downsides, the alternatives, and a practical framework you can use to know whether yours are ready.
Below, the points that landed hardest and the questions the conversation left open.
The conversation covered:
• Why is AI everywhere in finance now, but you can’t handle every decision?
• How does LLM-as-a-Judge work, why do people use it and what is it good for, where does it fall down?
• What are the alternatives, and what does each cost you?
• How do you know which kind of guardrail you should use?
• A guardrail readiness framework
Top takeaways
1) AI is everywhere in financial services, but so is the risk
The adoption of AI has moved faster than comprehension of how it works, and the areas like financial services where AI would be most useful are also the areas where getting it wrong carries the highest regulatory cost. Guardrails aren’t an optional layer in this environment; they’re the mechanism that determines whether AI is deployable at all.
“”Big AI labs will say that they do model safety testing. They will entrench guardrails into the training data but that is not a guardrail. You cannot separate it and test it. It’s difficult to give evidence.” — Alex Corsham”
2) In regulated work, consistency is the real problem, not accuracy
LLM-as-a-Judge is genuinely useful for the ambiguous, natural-language decisions where rules can’t reach. The same flexibility that makes it powerful also creates the questions worth designing around: its verdicts and explanations emerge from the same generation path, so the reasoning it produces is closer to a summary of the decision than a trace of it.
None of that means LLM-as-a-Judge doesn’t belong in a production system, it means it belongs as one voice among several, with the harness around it doing the work that keeps it consistent, auditable, and defensible.
“”The verdict and explanation are coming out in the same tangled path. It’s more of a post-hoc explanation, so for financial services, it’s very tempting, but you have to design for reproducibility.” — Alex Corsham”
3) Every guardrail has a cost, the question is which one you can afford to pay
The four options: human review, rule-based systems, fine-tuned classifiers, LLM-as-a-Judge, each have a natural home, and each carry a different kind of cost. Human review handles ambiguity best but doesn’t scale. Rules are cheap, fast, and auditable but struggle with the flexibility of natural language. Classifiers are reliable and specialised but need labelled data most teams don’t have. LLM-as-a-Judge works out of the box and handles nuance, but costs per call and needs supporting structure for consistency and audit.
“”One of the words we typically use to describe LLMs is ‘dynamic’ and ‘fluid.’ But actually they can become really brittle when you want to get really specific about how they should function.” — Dr Ailish McLaughlin”
4) The right guardrail strategy is an architectural one
Guardrail choice is really a risk-direction decision. Overblocking sends a customer to a contact centre; underblocking sends unauthorised advice into the market. Those failures aren’t the same size, and they demand different architectures. The EU AI Act sharpens this further: any deployed guardrail has to be defensible three years later, not just impressive in a demo.
“”Being a voice among several is absolutely a place for these kinds of things. They can be incredibly powerful but relying on just them as a single check is not the best approach.” — Alex Corsham”
5) A guardrail readiness framework
Before any guardrail goes near production, there are five questions worth asking of it:
“1. Can you demonstrate, reliably, that your guardrail is doing what you expect it to do before you deploy it? This covers things like having a suitable number of evaluation cases, covering not just the happy path cases, such that when you tweak prompts or components, you have confidence it is going to hold up 2. What is the risk of a failure in your guardrails? (similar to the last question) this governs how strenuous their guardrails actually need to be - is one check going to cut it? what happens in ambiguity? is automation even worth it here 3. Will it fit in your budget? Idea here being you could call some very expensive models to get what you need, but this might make the whole process totally worthless as an exercise 4. Where does this flow impact the user? For guardrails in chatbots, this shows up as a realtime check, adding 10s of seconds of latency for a complex check is unlikely to be acceptable 5. Will it hold up under audit? This covers explainability and whether it is just a reasoning trace from an LLM or something stronger, whether results of guardrails are being stored and can be retrieved later and will satisfy compliance, customers and the regulator”
6) The neurosymbolic approach to complex decisioning
Rather than asking a single model “is this a breach?”, UnlikelyAI ingests the policy into a symbolic structure and asks small, reliable questions instead. That lets you use smaller, cheaper and faster models, and, crucially, lets the system answer “we don’t know” on genuinely ambiguous cases and bias that answer towards failing closed. What you are left with is an explanation you can store and defend, not just a verdict you have to trust.
“”Rather than flip-flopping between two options, you have a classification for that ambiguous layer, and you can bias that in whatever direction you want.” — Alex Corsham”





