<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Evaluation on AI Science Report</title>
    <link>https://aiscience.uk/tags/evaluation/</link>
    <description>Recent content in Evaluation on AI Science Report</description>
    <generator>Hugo</generator>
    <language>en-us</language>
    <lastBuildDate>Wed, 15 Jul 2026 00:00:00 +0800</lastBuildDate>
    <atom:link href="https://aiscience.uk/tags/evaluation/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>When an AI Judge Gives an Unfair Score, the Bias Has a Shape You Can Touch</title>
      <link>https://aiscience.uk/posts/llm-judge-bias-mechanistic-interpretability-activation-steering/</link>
      <pubDate>Wed, 15 Jul 2026 00:00:00 +0800</pubDate>
      <guid>https://aiscience.uk/posts/llm-judge-bias-mechanistic-interpretability-activation-steering/</guid>
      <description>&lt;p&gt;If you ask ChatGPT to rate a response that begins with &amp;ldquo;GPT-4:&amp;rdquo; versus the same response labeled &amp;ldquo;GPT-2:&amp;rdquo;, you know what happens. The score drops. Not because the content changed (it didn&amp;rsquo;t) but because the label whispered something the model couldn&amp;rsquo;t ignore. This is LLM-as-judge bias, and until now, it&amp;rsquo;s been studied almost entirely from the outside: tweak the input, measure the score shift, repeat. A new paper from researchers at Alibaba, MBZUAI, USC, and Michigan asks a different question. When an LLM judge gives an unfair score, what&amp;rsquo;s happening inside the model?&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
