Last updated 3 weeks agoJob is active

Site Reliability Engineer

Keep the platform running through failures: component redundancy, backup storage, monitoring, and protection of the perimeter.

Responsibilities:

  • design redundancy for system components so a single failure does not stop game rounds;
  • own backup data storage and regularly verified restore procedures;
  • build monitoring, alerting, and incident response practices for high-load services;
  • operate DDoS protection and support fraud-protection systems around the platform perimeter.

Requirements:

  • experience operating high-load production systems in an SRE, DevOps, or platform role;
  • strong Linux, networking, and observability skills;
  • experience with incident management: detection, mitigation, honest postmortems;
  • automation-first mindset: infrastructure and recovery described as code.

Nice to have:

  • experience running JVM services in production;
  • experience with container orchestration or similar deployment stacks;
  • familiarity with security hardening and abuse mitigation.