As a Senior Site Reliability Engineer at Nebius, your primary purpose will be to ensure the reliability and performance of the inference platform, directly contributing to its scalability and cost-efficiency. Key responsibilities include designing telemetry pipelines, optimizing Kubernetes configurations, and implementing resilient infrastructure through Terraform. Required skills include deep knowledge of Kubernetes and experience with GPU workloads, complemented by scripting abilities in Pytho