Study flags 'linguistic illegibility' for LLM security
In a paper, James Mickens introduces the concept of linguistic illegibility: an LLM's externalized language and mechanistically probed linguistic features may not reflect its internal computation. Security mechanisms relying on linguistic self-reporting, such as chain-of-thought monitoring, constitutional self-critique and activation probing, can therefore never be fully sound; sandboxes need isolation techniques that don't depend on reading a model's linguistic state. The author calls taint tracking a promising approach and suggests robust virtualization and third-party auditing as complements.