Bo Li, Huan Sun, Somesh Jha, -Madison, Yevgeniy Vorobeychik, Chaowei Xiao, Xiaogeng Liu, Patrick Mcdaniel, Peiran Li, Edward Suh, Zhuoqing Mao
We lifted 3 functions out of this paper's own repositories and ran 3 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.
| Repository | Role | Ran |
|---|---|---|
| safolab-wisc/autodan-turbo | — | 3 of 3 |
| Function | Status | Where it lives |
|---|---|---|
| AutoDANTurbo | Ran | safolab-wisc/autodan-turbo/pipeline.py code served (permissive licence) · get_code("3b01bdc07291cb54") |
| Library | Ran | safolab-wisc/autodan-turbo/pipeline.py code served (permissive licence) · get_code("c8449f77c53c814a") |
| Log | Ran | safolab-wisc/autodan-turbo/pipeline.py code served (permissive licence) · get_code("aac277ca6868ec00") |
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
In this paper, we propose AutoDAN-Turbo, a black-box jailbreak method that can automatically discover as many jailbreak strategies as possible from scratch, without any human intervention or predefined scopes (e.g., specified candidate strategies), and use them for red-teaming. As a result, AutoDAN-Turbo can significantly outperform baseline methods, achieving a 74.3% higher average attack success rate on public benchmarks. Notably, AutoDAN-Turbo achieves an 88.5 attack success rate on GPT-4-1106-turbo. In addition, AutoDAN-Turbo is a unified framework that can incorporate existing human-designed jailbreak strategies in a plug-and-play manner. By integrating human-designed strategies, AutoDAN-Turbo can even achieve a higher attack success rate of 93.4 on GPT-4-1106-turbo.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2410.05295")
get_code_for_paper("2410.05295")
have("2410.05295")
Connect an agent — have() is free.