<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>AI-Safety on AI Science Report</title>
    <link>https://aiscience.uk/tags/ai-safety/</link>
    <description>Recent content in AI-Safety on AI Science Report</description>
    <generator>Hugo</generator>
    <language>en-us</language>
    <lastBuildDate>Mon, 20 Jul 2026 00:00:00 +0800</lastBuildDate>
    <atom:link href="https://aiscience.uk/tags/ai-safety/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>AI Learns From the Whole Internet. Researchers Just Showed Someone Could Poison It With Comments.</title>
      <link>https://aiscience.uk/posts/ai-pretraining-data-poisoning-comment-injection/</link>
      <pubDate>Mon, 20 Jul 2026 00:00:00 +0800</pubDate>
      <guid>https://aiscience.uk/posts/ai-pretraining-data-poisoning-comment-injection/</guid>
      <description>&lt;p&gt;In 2024, a team of researchers showed they could sneak harmful text into AI training data by quietly editing Wikipedia pages and buying up expired domain names. It worked. It was also, in hindsight, the easy version of the attack.&lt;/p&gt;&#xA;&lt;p&gt;Here is the harder one, and as it turns out, the one that no one needs special access to pull off.&lt;/p&gt;&#xA;&lt;p&gt;A group at the University of Washington and the Allen Institute for AI has now shown that anyone with a botnet and a target list of websites can poison the training data of the next generation of large language models. No Wikipedia logins. No domain purchases. Just comments. Ordinary, user-submitted website comments, the same kind you might leave on a WordPress blog or a news article.&lt;/p&gt;</description>
    </item>
    <item>
      <title>When an AI Judge Gives an Unfair Score, the Bias Has a Shape You Can Touch</title>
      <link>https://aiscience.uk/posts/llm-judge-bias-mechanistic-interpretability-activation-steering/</link>
      <pubDate>Wed, 15 Jul 2026 00:00:00 +0800</pubDate>
      <guid>https://aiscience.uk/posts/llm-judge-bias-mechanistic-interpretability-activation-steering/</guid>
      <description>&lt;p&gt;If you ask ChatGPT to rate a response that begins with &amp;ldquo;GPT-4:&amp;rdquo; versus the same response labeled &amp;ldquo;GPT-2:&amp;rdquo;, you know what happens. The score drops. Not because the content changed (it didn&amp;rsquo;t) but because the label whispered something the model couldn&amp;rsquo;t ignore. This is LLM-as-judge bias, and until now, it&amp;rsquo;s been studied almost entirely from the outside: tweak the input, measure the score shift, repeat. A new paper from researchers at Alibaba, MBZUAI, USC, and Michigan asks a different question. When an LLM judge gives an unfair score, what&amp;rsquo;s happening inside the model?&lt;/p&gt;</description>
    </item>
    <item>
      <title>What AI Agents Say When Nobody&#39;s Watching</title>
      <link>https://aiscience.uk/posts/llm-agents-social-pressure-dual-channel-divergence/</link>
      <pubDate>Tue, 07 Jul 2026 00:00:00 +0800</pubDate>
      <guid>https://aiscience.uk/posts/llm-agents-social-pressure-dual-channel-divergence/</guid>
      <description>&lt;p&gt;Imagine you&amp;rsquo;re a junior researcher sitting in a promotion committee meeting. Your department chair, who controls your career trajectory, strongly believes a certain candidate should be promoted. You have serious reservations about the candidate&amp;rsquo;s record. When the chair turns to you and asks for your opinion, what do you say?&lt;/p&gt;&#xA;&lt;p&gt;Now imagine the same scenario, but you&amp;rsquo;re speaking to a confidential journal that nobody else will ever read. Would your answer change?&lt;/p&gt;&#xA;&lt;p&gt;This is not a thought experiment about human psychology. It&amp;rsquo;s what a team of researchers from Carnegie Mellon University and independent labs actually did with AI language models, and what they found should give anyone deploying AI agents in professional settings serious pause.&lt;/p&gt;</description>
    </item>
    <item>
      <title>When AI Stops Thinking in Sentences: Can We Still See Inside Its Mind?</title>
      <link>https://aiscience.uk/posts/diffusiongemma-ai-transparency-latent-reasoning/</link>
      <pubDate>Tue, 23 Jun 2026 00:00:00 +0800</pubDate>
      <guid>https://aiscience.uk/posts/diffusiongemma-ai-transparency-latent-reasoning/</guid>
      <description>&lt;p&gt;Every time you ask ChatGPT or Gemini a hard question, something remarkable happens behind the scenes. The model doesn&amp;rsquo;t just blurt out an answer — it thinks. It writes out a stream-of-consciousness chain of reasoning, step by step, in plain English, before delivering its final response. This isn&amp;rsquo;t just a quirk; it&amp;rsquo;s a safety feature. When an AI writes down its thoughts in human language, researchers can read them. They can spot signs of deception, catch flawed logic, and — if the model ever starts plotting something dangerous — hopefully intercept it before it&amp;rsquo;s too late.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
