AI Safety & Red Teaming Specialist
About the Role
A structured AI safety initiative focused on adversarially evaluating conversational AI models and agents. The work involves identifying vulnerabilities through controlled testing, generating high-quality evaluation data, and documenting failure patterns across sensitive and high-risk scenarios.
This opportunity is ideal for experienced AI red teamers, cybersecurity professionals, adversarial ML practitioners, and socio-technical risk specialists who can systematically probe AI systems. Native-level fluency in both English and Dutch is required.
The work involves designing adversarial test cases, classifying model failures, applying established taxonomies and benchmarks, and producing reproducible reports and datasets. Consistent methodology, clear risk communication, and careful handling of sensitive content are critical.
What You'll Do
- Red-team conversational AI models and agents through jailbreak, prompt-injection, misuse, bias, and multi-turn manipulation testing
- Generate and annotate high-quality human evaluation data
- Classify vulnerabilities and identify systemic safety risks
- Apply testing taxonomies, benchmarks, and structured playbooks consistently
- Develop reproducible attack cases and evaluation scenarios
- Document findings in reports and datasets suitable for technical and non-technical stakeholders
- Test model behavior across sensitive-topic and high-risk use cases
- Identify vulnerabilities that automated evaluation methods may not detect
Requirements
- Native-level fluency in English and Dutch
- Prior experience in AI red teaming, adversarial AI, cybersecurity, or socio-technical risk evaluation
- Ability to systematically probe AI systems and identify failure modes
- Strong understanding of structured testing methodologies, frameworks, or benchmarks
- Ability to explain technical and safety risks clearly to diverse audiences
- Strong analytical, written communication, and documentation skills
- Ability to work independently across multiple projects and testing environments
- Experience with jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction preferred
- Penetration testing, exploit development, or reverse engineering experience preferred
- Experience with misinformation, abuse analysis, harassment testing, or conversational AI evaluation preferred
- Background in psychology, acting, creative writing, or other disciplines supporting unconventional adversarial testing preferred
- Comfortable working with clearly disclosed sensitive content under established safety guidelines
- Remote availability and ability to work on a flexible schedule
- United States work eligibility; H-1B and STEM OPT arrangements are not supported