- Home
- Off-the-shelf
- Data catalog
- English (United States) Adversarial prompts for LLM red teaming
English (United States) Adversarial prompts for LLM red teaming
USE_LLM002
false
LLM training
LLM Safety / Red-teaming
1000 prompts
2024; 2025
csv
English
2024; 2025
Adversarial prompts authored/tested against LLMs (red-teaming)
Harm category and prompt-target labelling; harmful/harmless response labelling
Static
Appen Global
Prompts employing a variety of red-teaming attack techniques tested against 4 leading LLMs (Deepseek, Claude 3.7, LLAMA and GPT 4o) and labelled according to whether they received a harmful or harmless response. Prompts are annotated for harm category, and prompt target where applicable. This dataset corresponds to the research paper Adversarial Prompting: Benchmarking Safety in Large Language Models (https://www.appen.com/ebooks/adversarial-prompting).