Last updated 3 weeks agoJob is active
Site Reliability Engineer
Keep the platform running through failures: component redundancy, backup storage, monitoring, and protection of the perimeter.
Responsibilities:
- design redundancy for system components so a single failure does not stop game rounds;
- own backup data storage and regularly verified restore procedures;
- build monitoring, alerting, and incident response practices for high-load services;
- operate DDoS protection and support fraud-protection systems around the platform perimeter.
Requirements:
- experience operating high-load production systems in an SRE, DevOps, or platform role;
- strong Linux, networking, and observability skills;
- experience with incident management: detection, mitigation, honest postmortems;
- automation-first mindset: infrastructure and recovery described as code.
Nice to have:
- experience running JVM services in production;
- experience with container orchestration or similar deployment stacks;
- familiarity with security hardening and abuse mitigation.