DevOps Engineer – Cloud Infrastructure & AI Evaluation
About the Role
A structured technical evaluation initiative focused on improving advanced AI systems through realistic DevOps and cloud infrastructure scenarios. The work applies practical infrastructure engineering expertise to production-style tasks involving automation, containers, orchestration, infrastructure as code, monitoring, and incident response.
This opportunity is ideal for DevOps, SRE, Platform, and Infrastructure Engineers with hands-on experience building and maintaining reliable cloud environments. Strong programming ability, Kubernetes, CI/CD, Terraform, Linux, and cloud platform expertise are central to the role; prior AI experience is not required.
The work involves creating reproducible engineering tasks, developing reference implementations, reviewing AI-generated solutions, and documenting precise corrections. Technical correctness, reliability, reproducibility, and clear written reasoning are critical.
What You'll Do
- Design realistic DevOps tasks based on production-style infrastructure scenarios.
- Create challenges involving CI/CD, containers, Kubernetes, infrastructure as code, monitoring, and incident response.
- Develop working reference solutions with clear setup, execution, and validation procedures.
- Review AI-generated infrastructure and engineering outputs for correctness and reliability.
- Identify implementation failures and document precise technical corrections with supporting reasoning.
- Validate task environments for reproducibility and consistent outcomes.
- Maintain rigorous quality standards across infrastructure tasks and reference solutions.
- Provide structured feedback to improve future task design and evaluation criteria.
- Apply infrastructure automation and troubleshooting expertise to complex technical scenarios.
- Collaborate with technical contributors in a remote project environment.
Requirements
- 3+ years of professional experience in DevOps, SRE, Platform Engineering, or Infrastructure Engineering.
- Strong programming ability in Python, Go, Bash, or TypeScript, including production-oriented code beyond basic scripting.
- Practical experience with CI/CD platforms such as GitHub Actions, GitLab CI, or Jenkins.
- Hands-on experience with Docker and Kubernetes.
- Experience with at least one major cloud platform, such as AWS, GCP, or Azure.
- Strong Infrastructure as Code experience with Terraform, Ansible, Pulumi, or CloudFormation.
- Advanced Linux administration and troubleshooting skills.
- Strong Git proficiency and modern software development workflow knowledge.
- Professional written and spoken English.
- Availability for at least 20 hours per week.
- Strong attention to technical detail, reproducibility, and correctness.
- Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry is preferred.
- Background in security, secrets management, or platform reliability is preferred.
- Open-source contributions or technical writing experience is preferred.
- Previous AI training-data creation or model evaluation experience is preferred.