We haven't found a way to prevent AI chat systems from unintentionally providing dangerous instructions when questions are phrased indirectly.
open
Global / Unspecified, Global
WS00860
Safety filters designed to block clearly harmful requests can sometimes be bypassed through indirect or creatively phrased questions that still elicit dangerous information. Closing this gap while preserving the system's general usefulness remains an unresolved balancing act.