Company Description:
Sutherland is seeking a reliable person to join us as Platform and Cloud Engineer will be responsible to bridges the gap between IT operations and cybersecurity, focusing on the continuous protection of infrastructure and data through automation, monitoring, and proactive defense.
Job Description:
KEY RESPONSIBILITIES
Platform Architecture & Strategy (Senior-focused)
- Lead the architecture and deployment of complex multi-cloud solutions (GCP/AWS), including networking, compute, storage, and multi-environment design.
- Define platform standards and reusable patterns (e.g., Terraform modules, cluster blueprints).
- Evaluate emerging cloud technologies to enhance cloud strategy and roadmap.
- Drive cost optimization and FinOps practices, including right-sizing and governance.
- Own platform reliability, scalability, capacity planning, and disaster recovery design.
Infrastructure Provisioning & Automation
- Deploy, configure, and manage cloud infrastructure across GCP and AWS.
- Write and maintain infrastructure as code (Terraform primary; CloudFormation/ARM where applicable).
- Maintain multi-environment infrastructure consistency (dev, staging, prod).
- Automate provisioning, configuration, and operational tasks to reduce manual toil.
Kubernetes & Container Platform
- Design, operate, and support production-grade Kubernetes clusters (GKE preferred).
- Manage upgrades, autoscaling, node pools, namespaces, and RBAC policies.
- Own Helm/Kustomize standards for application deployment.
- Support or manage service mesh (Istio) for traffic management, mTLS, observability, and security.
- Define and promote golden paths for safe and consistent app deployment.
CI/CD, Monitoring & Operational Support
- Build and maintain GitLab CI/CD pipelines for infrastructure and application delivery.
- Monitor platform health using Datadog — metrics, logs, traces, SLOs, and alerts.
- Troubleshoot issues, support incident response, and lead post-incident reviews (senior).
- Maintain runbooks and dashboards; lead or support on-call rotations.
Leadership & Collaboration (Senior-focused)
- Mentor junior and mid-level engineers, set technical direction, and review designs.
- Collaborate effectively with global, cross-functional, and on/offshore teams.
- Communicate architecture decisions clearly to technical and business stakeholders.
TECH STACK
Required:
- Cloud Platforms: GCP (Compute Engine, GKE, VPC, Storage, IAM, Load Balancing), AWS (EC2, EKS, VPC, S3, IAM)
- Kubernetes: GKE architecture, upgrades, Helm/Kustomize, autoscaling, RBAC
- Infrastructure as Code: Terraform (multi-environment, reusable modules, remote state)
- CI/CD: GitLab pipeline development and support
- Monitoring: Datadog (metrics, logs, traces, alerting, SLOs)
- Service Mesh: Istio (traffic management, mTLS, observability)
Good to have:
- GitOps tools (ArgoCD, Flux)
- Cloud FinOps tooling and cost optimization experience
- Scripting languages (Python, Go, Bash)
- Vault, Packer, service catalogs, and self-service platforms
- Experience in regulated environments (HIPAA, SOC 2, ISO 27001)
Qualifications:
REQUIREMENTS
Must have:
- Platform Engineer: 3+ years in cloud infrastructure/platform or DevOps engineering.
- Senior Platform Engineer: 8+ years, with leadership or architect-level responsibilities.
- Deep, hands-on expertise with GCP and/or AWS cloud platforms.
- Strong Kubernetes production experience, preferably GKE.
- Expert-level Terraform skills with reusable modules and multi-env IaC.
- Proven CI/CD automation experience at scale.
- Ability to troubleshoot complex cloud environment issues.
- Collaborative teamwork mindset with clear communication skills.
Nice to have:
- Certifications: Google Professional Cloud Architect, AWS Solutions Architect Professional, Certified Kubernetes Administrator (CKA).
- Experience with GitOps and progressive delivery techniques.
- Prior FinOps or cloud cost optimization role experience.
- Scripting for automation and Linux system administration background.
HOW SUCCESS IS MEASURED
- Platform reliability: uptime and SLO achievement for shared services.
- Automation coverage and manual toil reduction.
- Adoption and enforcement of platform standards and self-service deployment paths.
- Cloud cost efficiency achieved through optimization.
- Timely and successful delivery of platform initiatives and projects.
- Mean time to recovery (MTTR) for platform-impacting incidents.
Additional Information:
All your information will be kept confidential according to EEO guidelines.