We are looking for a Site Reliability Engineer who can design, build, and operate resilient infrastructure and CI/CD pipelines across multiple cloud platforms (AWS, Azure, GCP). This role is central to enabling reliable, scalable, and secure deployment of applications and AI/data platforms, with a strong focus on automation, observability, and infrastructure-as-code across heterogeneous cloud environments.
Site Reliability Engineer (Multi-Cloud Infrastructure & Pipelines)
Role Title: Site Reliability Engineer (SRE) – Multi-Cloud Infrastructure & CI/CD
Domain: Cloud Infrastructure, DevOps/SRE, Platform Engineering
Responsibilities
Key Responsibilities
- Design, build, and maintain cloud-agnostic infrastructure using Infrastructure-as-Code (Terraform, Pulumi, or equivalent) across AWS, Azure, and GCP
- Build and manage CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Azure DevOps, Cloud Build) supporting multi-cloud deployments
- Define and implement SRE practices: SLIs/SLOs/SLAs, error budgets, incident response, and on-call processes
- Set up monitoring, logging, and observability stacks (Prometheus, Grafana, Datadog, Cloud Monitoring/Stackdriver, ELK/EFK) across environments
- Architect and manage container orchestration platforms (Kubernetes — EKS/AKS/GKE) and containerization (Docker) for portable, cloud-agnostic workloads
- Automate provisioning, scaling, patching, and configuration management (Ansible, Chef, or Puppet as needed)
- Implement disaster recovery, backup, high-availability, and multi-region/multi-cloud failover strategies
- Drive cost optimization and capacity planning across cloud providers
- Build self-healing systems and automate incident remediation to reduce toil
- Conduct root cause analysis (RCA) and post-incident reviews; drive reliability improvements
- Implement security best practices: IAM policies, secrets management (Vault, Cloud KMS), network security, and compliance guardrails across clouds
- Collaborate with development teams to embed reliability, observability, and deployment best practices into the software delivery lifecycle (DevSecOps/shift-left)
- Support migration or workload portability initiatives between cloud providers or hybrid/on-prem environments
Requirements
Required Skills & Experience
- 5+ years of experience in SRE, DevOps, or Infrastructure Engineering roles
- Proven hands-on experience across at least two of the three major cloud providers (AWS, Azure, GCP) — true multi-cloud experience strongly preferred
- Strong expertise in Infrastructure-as-Code (Terraform mandatory; CloudFormation/Bicep/Pulumi a plus)
- Deep knowledge of Kubernetes and container orchestration in production environments
- Strong scripting/programming skills (Python, Go, or Bash)
- Experience building and maintaining CI/CD pipelines end-to-end (build, test, deploy, rollback)
- Solid understanding of networking fundamentals (VPCs, load balancers, DNS, service mesh — Istio/Linkerd a plus)
- Experience with observability and monitoring tooling, and setting up alerting that ties to meaningful SLOs
- Familiarity with GitOps practices (ArgoCD, FluxCD)
- Strong incident management experience — triage, escalation, blameless postmortems
- Understanding of security and compliance requirements in cloud environments (encryption, IAM, network segmentation)
Preferred Qualifications
- Cloud certifications across multiple providers (e.g., AWS Solutions Architect/SysOps, Azure Administrator/DevOps Engineer, GCP Professional Cloud DevOps Engineer)
- Certified Kubernetes Administrator (CKA) or CKAD
- Experience with service mesh, API gateways, and edge/CDN configuration
- Experience supporting AI/ML or data platform infrastructure (GPU provisioning, vector databases, model serving infra)
- Prior experience in telecom, enterprise SaaS, or large-scale regulated environments
- Familiarity with FinOps practices for multi-cloud cost governance