Anaplan

Principal Platform Engineer (High Availability & Disaster Recovery)

Hiring

Anaplan

Gurugram · IndiaFull-time· Today

Experience

8+ years

Work Type

Full-time

Domain

Cloud, SaaS

Leadership

Required

Core SkillsMust Have
HA engineeringreal-time data replicationdatabase clusteringautomated failoverTerraformAnsibleLinux internalsKubernetesBGP routingDNS managementAnycastCDNshybrid networkingcontainer orchestration
What You'll Do
1Drive the HA Strategy & Roadmap: own end-to-end design and technical execution of global High Availability roadmap ensuring continuous platform uptime across cloud and on-premises environments
2Build & implement HA architectures: code, configure, and engineer active-active clustering, global load balancing, and real-time database replication to eliminate single points of failure
3Establish the Resiliency Practice: build and champion foundational Resiliency Engineering framework from scratch, defining platform-wide standards for fault tolerance and self-healing
4Influence engineering teams: partner with cross-functional engineering squads to embed self-healing mechanisms and automated recovery protocols
5Bootstrap Chaos Engineering: introduce and champion first-ever Chaos Engineering program, define strategy, establish safe guardrails, and prepare organization to execute automated fault-injection and live-fire failover drills
6Optimize uptime & manage risks: monitor infrastructure performance KPIs, mitigate shortage or overload risks, and ensure compliance with strict customer SLA requirements
Nice to HaveOptional
SaaS products supportIncident ManagementPost MortemsobservabilitymonitoringAWSGCPAzureconfiguration managementinfrastructure as code
Domain
CloudSaaS
LeadershipMentoring or team leadership preferred
Principal Platform Engineer (High Availability & Disaster Recovery) at Anaplan | Switchly