LLM

When an AI Judge Gives an Unfair Score, the Bias Has a Shape You Can Touch

When an AI Judge Gives an Unfair Score, the Bias Has a Shape You Can Touch

If you ask ChatGPT to rate a response that begins with “GPT-4:” versus the same response labeled “GPT-2:”, you know what happens. The score drops. Not because the content changed (it didn’t) but because the label whispered something the model couldn’t ignore. This is LLM-as-judge bias, and until now, it’s been studied almost entirely from the outside: tweak the input, measure the score shift, repeat. A new paper from researchers at Alibaba, MBZUAI, USC, and Michigan asks a different question. When an LLM judge gives an unfair score, what’s happening inside the model?

What AI Agents Say When Nobody's Watching

What AI Agents Say When Nobody's Watching

Imagine you’re a junior researcher sitting in a promotion committee meeting. Your department chair, who controls your career trajectory, strongly believes a certain candidate should be promoted. You have serious reservations about the candidate’s record. When the chair turns to you and asks for your opinion, what do you say?

Now imagine the same scenario, but you’re speaking to a confidential journal that nobody else will ever read. Would your answer change?

This is not a thought experiment about human psychology. It’s what a team of researchers from Carnegie Mellon University and independent labs actually did with AI language models, and what they found should give anyone deploying AI agents in professional settings serious pause.

Citations Make AI Hallucinate More, Not Less — A Major New Study

Citations Make AI Hallucinate More, Not Less — A Major New Study

You’d think that adding a citation to an AI’s answer would make it more trustworthy. A link to a scientific paper. A reference to a legal precedent. A footnote pointing to a medical journal. Surely the AI has checked its sources, right?

Wrong. A landmark new study accepted at ICML 2026 — one of the world’s most prestigious AI conferences — reveals something deeply counterintuitive: citations make large language models more likely to hallucinate, not less.