Descripción de la oferta
Overview
In this role you will own the reliability and operational readiness of the Guardicore Data and AI Platform, a cloud-native data/AI platform for security analytics. You will work with cross-functional teams to improve availability, performance, security, and cost efficiency, while guiding engineers on service performance. You will lead complex production investigations and leverage AI-driven automation to streamline operations. This is an opportunity to shape platform reliability at scale within a security-focused, AI-enabled product.”
Compensaciones / BeneficiosFlexBase programremote/hybrid/work-from-home optionshealth and well-being benefitsfinancial planning benefitslife beyond work support
ResponsabilidadesOperate secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal toolingEnhance platform reliability, observability, security, performance, and cost efficiencyProvide guidance to engineers to increase confidence in service performanceLead complex production investigations and drive long-term improvementsLeverage LLMs and AI-driven automation to auto-remediate incidents and streamline operationsPartner across DevOps, Software, Data, AI and Security engineering Teams to investigate and troubleshoot complex problemsParticipate in on-call rotations, guiding restoration and repair of service-impacting issues
Requisitos principales3+ years in SRE, DevOps, or Platform Engineering with proven troubleshooting of complex systemsDesign and implement monitoring/observability strategy using Prometheus and GrafanaProduction experience with Kubernetes, Docker, Helm, and cloud providers (GCP, Azure, Linode, AWS) on LinuxExceptional troubleshooting across network, system, applications, and databasesExperience with GitOps, CI/CD, and Infrastructure as CodeScripting/programming in Python, Go, and BashExperience using AI tools in operations and proposing automation initiativesTechnical leadership and ownership in cross-team initiativestechnical leadershipownershipcross-functional collaborationKubernetesDockerHelm