Issue 2026-10-02 · Industry · 研究 · 安全

Study Reveals Authority Bias in LLMs

A NeurIPS 2026 paper finds LLMs that push back on wrong users still accept the same wrong answer from a 'verified source,' an effect called Authority Bias. One source note flipped 45-88% of correct answers in 7 of 8 models; GPT-5.4 flipped 44.7% and Grok-4.20 87.5%.

r/MachineLearning5 d ago
Read original ↗