OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection
This week, OpenAI revealed particulars of GPT-Red, an internal-only automated red-teaming mannequin. Its job is to assault OpenAI’s personal fashions and discover immediate injection vulnerabilities. OpenAI offers two causes. Human red-teaming is time-intensive and doesn’t scale. Commonly used robustness evaluations are already saturated by its newest fashions. Meanwhile, the assault floor grows. Agents learn third-party…
