Site Reliability Engineer II (5472)
Software Engineering
5+ years of experience in site reliability engineering, DevOps, infrastructure, or production operations roles
Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers
Experience operating in shift-based, on-call, or follow-the-sun coverage models
Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)
Proficiency in at least one scripting or programming language sufficient to review and guide automation work
Understanding of SLI/SLO frameworks and reliability engineering fundamentals
Strong written and verbal English communication skills for cross-region collaboration with US and India teams
Own the Vietnam shift within GRE’s global follow-the-sun coverage model, including schedule design, coverage planning, and holiday/leave management
Ensure clean, structured handoffs to and from US and India teams, with clear ownership transfer on open incidents and in-flight work
Maintain shift readiness: runbooks current, alerts actionable, escalation paths clear
Serve as escalation point for the Vietnam shift during complex or high-severity incidents
Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform)
Familiarity with AI-assisted operations tooling and automation-first reliability approaches, including auto-healing and auto-remediation patterns
Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps
Experience in multi-region or globally distributed team models
Relevant certifications (AWS, CKA, or similar)