The Senior Site Reliability Engineer at Nebius will play a critical role in ensuring the reliability and performance of the inference platform that supports various AI models. Key responsibilities include designing telemetry pipelines, tuning Kubernetes for GPU efficiency, and creating Terraform modules to enhance cluster resilience. Candidates should have substantial experience with Kubernetes and GPU-heavy workloads, alongside strong scripting skills in Python or Bash.