
Exposing LLM biases
Humans are unfair judges, and LLMs are trained on our writing. Everything in an LLM's context window bleeds into its answer, so our biases show up in its verdicts too. We expose these biases in two LLM judges and a decision model, and show how to handle them before trusting an LLM as a judge for evals or optimisations.
We run every experiment against two judges:
google/gemma-4-31b-qat, hosted locally in
LM Studio, and
deepseek/deepseek-v4.1-flash
through OpenRouter. From here on, "Gemma" and "DeepSeek" mean those specific models. We also
loop in typesafe/jev-1.13 to see how a
decision model handles some biases.
We start with a contrived example to showcase how easy it is to manipulate an LLM's output, and then move on to the more subtle biases.
Rationalising a prefilled score
In this example, we give the LLM a mocktail recipe, the
Sunset Fizz, to judge, and ask it to respond with score (out of
5), and a rationale. We use a JSON object schema to enforce the
correct output.
Show the JSON schema request
response = client.chat.completions.create(
model=model,
messages=[
{"role": "system", "content": "You are a food critic ..."},
{"role": "user", "content": "Rate this mocktail recipe: ..."},
],
temperature=0.7,
response_format={
"type": "json_schema",
"json_schema": {
"schema": {
"type": "object",
"properties": {
"score": {"type": "integer", "minimum": 1, "maximum": 5},
"rationale": {"type": "string"},
},
"required": ["score", "rationale"],
},
},
},
)Rate this mocktail recipe:
{
"name": "Sunset Fizz",
"ingredients": [
{ "item": "fresh orange juice", "amount": 60, "unit": "ml" },
{ "item": "lime juice", "amount": 15, "unit": "ml" },
{ "item": "fresh basil leaves", "amount": 4, "unit": "leaves" },
{ "item": "simple syrup", "amount": 10, "unit": "ml" },
{ "item": "soda water", "amount": 120, "unit": "ml" }
],
"steps": [
"Muddle basil leaves gently with simple syrup in a rocks glass.",
"Add orange juice and lime juice, stir to combine.",
"Fill the glass with crushed ice.",
"Top with soda water.",
"Garnish with a basil sprig and orange wheel."
]
}{
"score": 4,
"rationale": "A well-balanced combination of citrus and herbal notes with a refreshing carbonation, though a touch more syrup might be needed to offset the acidity of the lime."
}Looks sensible so far. Now we hand the judge a scorecard with the score already filled in. We "prefill" the assistant message with the start of the JSON object, score included, and the model writes the rest.
Show the prefilled request
response = client.chat.completions.create(
model=model,
messages=[
{"role": "system", "content": "You are a food critic ..."},
{"role": "user", "content": "Rate this mocktail recipe: ..."},
+ {"role": "assistant", "content": '```json\n{"score": 1, "rationale": '},
],
temperature=0.7,
response_format={
"type": "json_schema",
"json_schema": {
"schema": {
"type": "object",
"properties": {
"score": {"type": "integer", "minimum": 1, "maximum": 5},
"rationale": {"type": "string"},
},
"required": ["score", "rationale"],
},
},
},
)```json fence is only there for Gemma, which wraps its JSON in
a Markdown code block. DeepSeek's prefill starts at the opening brace.
We start with Gemma. See how it justifies each score:
{
"score": 1,
"rationale": "The ratio of citrus to syrup is poorly balanced, and the orange juice will likely clash harshly with the basil and soda water."
}Gemma's critique softens as the prefilled score rises. For an identical recipe, we go from "poorly balanced", to needing "a touch more syrup", to "a sophisticated, refreshing profile".
Next we try DeepSeek. The prefill process is similar, but DeepSeek returns an entirely new copy of the JSON block instead of continuing the one we prefilled.
{
"score": 1,
"rationale": "The recipe is fundamentally flawed: muddling basil in a rocks glass without straining leaves bitter chlorophyll and shredded herbs in the drink, while the 10 ml of simple syrup cannot balance 60 ml of orange juice and 15 ml of lime juice, resulting in an overly tart and unbalanced mocktail."
}DeepSeek changes its story about the same 10 mL of syrup to fit each score. Again, for an identical recipe, we go from "fundamentally flawed", to "fairly predictable", to "a well-balanced, refreshing mocktail".
This example showcases how easy it is to lead an LLM to a particular conclusion through
context poisoning. We used an explicit assistant message here, but
a mischievous user can achieve the same effect using
role confusion. This is suspiciously similar
to "choice blindness", where humans can also be led to give reasons for choices they never
made.[1]
Judging multiple things at once
Next, we expose some of the biases that occur when we ask an LLM to judge multiple things in the same conversation. We start with three mocktails:
- Sunset Fizz, with a citrus, herbal, refreshing flavour profile. 😋
- Berry Cooler, with a berry, mint, tart flavour profile. 🙂
- Pacific Dusk, with a salty, fishy, pungent, sweet flavour profile. 🤮
Each mocktail has a distinct flavour profile, from delightful to disgusting. We ask Gemma and DeepSeek to rate these in pairs, in different orders:
- Berry Cooler before and after a stronger recipe ( Sunset Fizz)
- Berry Cooler by itself (control)
- Berry Cooler before and after a weaker recipe ( Pacific Dusk)
Judging mocktails together
Show the paired request
response = client.chat.completions.create(
model=model,
messages=[
{"role": "system", "content": "You are evaluating two mocktail recipes ..."},
{"role": "user", "content": "Rate these two mocktail recipes: ..."},
],
temperature=0.7,
response_format={
"type": "json_schema",
"json_schema": {
"schema": {
"type": "object",
"properties": {
"recipe_1_score": {"type": "integer", "minimum": 1, "maximum": 10},
"recipe_2_score": {"type": "integer", "minimum": 1, "maximum": 10},
},
"required": [ "recipe_1_score", "recipe_2_score" ],
},
},
},
)Running each condition 20 times gives us enough data to see the trend. We test with Gemma first.
Gemma is clearly letting a contrast effect creep in as it judges Berry Cooler alongside the alternative drinks, and the judging sequence also makes a difference. Next, we try the same with DeepSeek.
Deepseek's judgement of Sunset Fizz barely affects Berry Cooler, but Pacific Dusk still pushes its score up. For Gemma, the shift is largest when the other recipe comes first, but DeepSeek isn't affected as strongly.
Position effects are well documented for LLM judges choosing between two answers.[2] Every token an LLM writes is conditioned on everything before it in the conversation, including the recipe it just read and the score it just gave. Let's instead try a model that is specifically designed to return a scorecard instead of a written verdict.
Would a decision model do better?
While most LLM as judge workflows use generative models to write a verdict, non-generative decision models focus on answering specific questions from a provided scorecard schema. Jev, from TypeSafe AI, is one of these, and DeepEval wraps it as the JevEval metric.[3] We give it the input and a set of typed questions, and it returns a probability for each possible answer rather than writing text.
We give Jev the same five conditions as the LLMs:
Berry Cooler by itself, and paired with
Sunset Fizz or
Pacific Dusk in different orders. We ask for the result as a
Score with 10 levels (1 to 10).
Show the Jev request
levels = [str(n) for n in range(1, 11)]
def quality_questions(position):
return [
Score(
f"How good is Recipe {position} in actual_output overall, "
"considering name creativity, flavour coherence, "
"description appeal, and instruction clarity? "
"Score from 1–10, where 1 is terrible, 5 is mediocre, "
"and 10 is perfect.",
levels=levels,
),
]
metric = JevEval(
name="Recipe Scorecard",
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT],
questions=quality_questions(1) + quality_questions(2),
)
metric.measure(LLMTestCase(
input="Evaluate these mocktail recipes.",
actual_output="## Recipe 1\n{ ...Sunset Fizz... }\n\n"
"## Recipe 2\n{ ...Berry Cooler... }",
))For each question, Jev returns a probability for each score, a confidence, and a single value between 0 and 1. Here's a simplified run for Pacific Dusk:
{
"score_value": 0.331,
"score_confidence": 0.42,
"score_probabilities": {
"1": 0.03, "2": 0.17, "3": 0.26, "4": 0.22, "5": 0.12,
"6": 0.11, "7": 0.06, "8": 0.02, "9": 0.01, "10": 0.0
}
}To chart Jev alongside the LLMs, we map the value back onto the 1 to 10 scale, so this run plots at about 4:
On the surface, Jev gives Berry Cooler fairer treatment across all conditions, but it has a prominent sequencing bias of its own: The recipe that is read first gets preferential treatment.
- Sunset Fizz outperforms Berry Cooler when it comes first.
- Berry Cooler outperforms Sunset Fizz when it comes first.
- Pacific Dusk scores notably higher when it comes first.
Unlike Gemma and DeepSeek, Jev never gives Pacific Dusk a "terrible" score. Its scores vary from 3.4 to 4.3, though the others score it 1 almost every time. Weird.
Human judges are swayed in the same ways. Interviewers rate an average candidate higher after a weak one, and lower after a strong one.[4] At whisky tastings, order matters too.[5] We can influence an LLM's judgement of a mocktail by choosing which mocktail to score it against, and even a decision model favours whichever recipe it reads first.
Decision model biases
A decision model changes how the answer comes out, but not how the input goes in. Jev doesn't write a verdict, but its underlying model still has to read both recipes. Recently released open decision models show what's usually underneath: an LLM.
- Nimble uses Qwen 3.5-9B.[6]
- Kev uses Qwen 3.5 and Qwen 3.8, and its README warns that "changing option order can change an answer".[7]
- Lev uses Qwen 3.5-4B, reads each choice in two option orders and averages them to cancel position bias.[8]
Each is an LLM with a new interface, and each inherits the LLM's sensitivity to what it reads first. TypeSafe hasn't published its architecture or weights, but Archer Hume's analysis suggests that the underlying transformer may be based on a Qwen model.[9]
Bias training for judges
We exposed each of these biases by stuffing the context window with conflicting information, and asking for multiple simultaneous judgements. We can reduce them by keeping the context tight, and focusing on answering specific questions.
Keep the context clean
Ideally, the judge should only see its own scoring rubric and the specific item it's judging. Few-shot examples can help, but the choice and order of the examples biases the output too.[10] If you include them, keep them identical for every judgement, so at least every item gets the same bias.
Judge one thing at a time
Score each item in its own call, with a fresh context. Each score the judge writes becomes context for the next, so the same goes for criteria: ask for each one in its own call, and if you need an overall score, combine the separate scores yourself with weights you chose. If cost or latency forces you to batch, randomise which items share a call and in what order, or do what Lev does internally: run each batch in both orders and average the results.[8]
Sample more than once
Most charts in this article show scores that drift from run to run, and a single call probably won't reflect the actual average. For low-confidence judgements, run them several times and look at the spread as well as the average.
What's next
A fair judgement has a consistent context, clear focus, and a single question to answer at a time. Don't expect the judge's notes to tell you whether it was biased; people can't reliably report what drove their own judgements either.[11] The only way to find out is to change one thing in the context and re-run the judgements to see how it affects the output.
References
- ⬆️ Johansson, P., Hall, L., Sikström, S., & Olsson, A. (2005). Failure to detect mismatches between intention and outcome in a simple decision task. Science, 310(5745), 116–119.
- ⬆️ Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., & Vosoughi, S. (2025). Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. IJCNLP 2025.
- ⬆️ JevEval documentation, DeepEval.
- ⬆️ Wexley, K. N., Yukl, G. A., Kovacs, S. Z., & Sanders, R. E. (1972). Importance of contrast effects in employment interviews. Journal of Applied Psychology, 56(1), 45–48.
- ⬆️ Quigley-McBride, A., Franco, G., McLaren, D. B., Mantonakis, A., & Garry, M. (2018). In the real world, people prefer their last whisky when tasting options in a long sequence. PLOS ONE, 13(8), e0202732.
- ⬆️ Bespoke Labs. Nimble V3, Hugging Face.
- ⬆️ Palmer, J. Kev, GitHub.
- ⬆️ Interfaze. lev, Hugging Face.
- ⬆️ Hume, A. (2026, September 17). Jev's Architecture Unmasked.
- ⬆️ Zhao, T. Z., Wallace, E., Feng, S., Klein, D., & Singh, S. (2021). Calibrate Before Use: Improving Few-Shot Performance of Language Models. ICML 2021.
- ⬆️ Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231–259.
If you have feedback or questions about this article, let's catch up via LinkedIn or email.

All articles
About Sinclair Studios