📰 News LLM Safety Training Compresses Risk Rather Than Removing It — New Research Warns of Fragile Guardrails
Safety training doesn't remove harmful capabilities from AI models — it packs them tighter. New mechanistic interpretability research reveals why current alignment approaches are fundamentally brittle.