Jun Yu, Yu Li, Min Zhang, Jing Li, Daojing He, Wenya Wang, Weiyang Guo
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
The proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures. However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses. In this paper, we propose the Multi-Turn Safety Alignment (MTSA) framework, to address the challenge of securing LLMs in multi-round interactions. It consists of two stages: In the thought-guided attack learning stage, the redteam model learns about thought-guided multiround jailbreak attacks to generate adversarial prompts. In the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction. Furthermore, we introduce a multi-turn reinforcement learning algorithm based on future rewards to enhance the robustness of safety alignment. Experimental results show that the red-team model exhibits state-of-the-art attack capabilities, while the target model significantly improves its performance on safety benchmarks. Code is available at https://github.com/yuki-younai/MTSA WARNING: This paper contains potentially offensive and harmful text. ′ H+1 ) of the ending state.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2505.17147")
get_code_for_paper("2505.17147")
have("2505.17147")
Connect an agent — have() is free.