We haven't found a way to prevent AI chat systems from unintentionally providing dangerous instructions when questions are phrased indirectly.
open
Global / Unspecified, Global
Safety filters designed to block clearly harmful requests can sometimes be bypassed through indirect or creatively phrased questions that still elicit dangerous information. Closing this gap while preserving the system's general usefulness remains an unresolved balancing act.
Citation ID: WS00860
Title: We haven't found a way to prevent AI chat systems from unintentionally providing dangerous instructions when questions are phrased indirectly.
URL: https://worldsolve.org/index.php?api=problem&id=860