Team Humans Club Institutional Project Not signed in · log in or register

World Solve

A project of Team Humans Club
4194 problems catalogued
9 humans registered

We haven't found a way to prevent AI chat systems from unintentionally providing dangerous instructions when questions are phrased indirectly.

open Global / Unspecified, Global WS00860
Safety filters designed to block clearly harmful requests can sometimes be bypassed through indirect or creatively phrased questions that still elicit dangerous information. Closing this gap while preserving the system's general usefulness remains an unresolved balancing act.
Created at: 2026-07-23T19:26:59Z
Click to copy citation
WS00860 | World Solve | https://worldsolve.org/index.php?view=problem&id=860

Related Problems