When an AI Judge Gives an Unfair Score, the Bias Has a Shape You Can Touch
If you ask ChatGPT to rate a response that begins with “GPT-4:” versus the same response labeled “GPT-2:”, you know what happens. The score drops. Not because the content changed (it didn’t) but because the label whispered something the model couldn’t ignore. This is LLM-as-judge bias, and until now, it’s been studied almost entirely from the outside: tweak the input, measure the score shift, repeat. A new paper from researchers at Alibaba, MBZUAI, USC, and Michigan asks a different question. When an LLM judge gives an unfair score, what’s happening inside the model?