Senior Platform Engineer - Cloud Infrastructure
Listed on 2026-08-04
-
Software Development
DevOps, Cloud Engineer - Software
About Salesforce
Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn't a buzzword - it's a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.
Ready to level-up your career at the company leading workforce transformation in the agentic era? You're in the right place! Agentforce is the future of AI, and you are the future of Salesforce.
Platform Engineering - Cloud InfrastructureOverview of the Role
SMTS role is part of our Platform Engineering team within the Cloud Infrastructure organization. Platform Engineering is made up of platform engineers, SREs, and Dev Ops specialists who design, build, and operate the internal developer platform powering hundreds of Kubernetes clusters across AWS, Azure, GCP, and OCI. Whether we are automating cluster lifecycle management, hardening our Git Ops delivery pipelines, or embedding AI into our daily engineering and infrastructure operations workflows, we strive to give every product team a fast, secure, and reliable path to production.
We are looking for a Senior Member of Technical Staff to accelerate the evolution of our multi-cloud platform and operational resilience. In addition to building and operating core platform services in Go and Python, you'll get a chance to shape our Git Ops, continuous deployment, and automated operations strategy, drive automation at fleet scale, and act as an AI amplifier for the team - multiplying engineering and operational output by championing agentic, AI-assisted engineering and automated remediation practices in everything we build and operate.
WhatYou'll Actually Be Doing
Success will be measured by the reliability, scalability, operational resilience, and developer experience of the platform services you build and operate across our multi-cloud Kubernetes fleet.
Design, build, and operate platform services and infrastructure automation in Go and Python that manage Kubernetes clusters at scale across AWS, Azure, GCP, and OCI.
Build and improve continuous deployment pipelines using Git Ops tooling (Flux, Argo CD) and infrastructure-as-code frameworks (Pulumi, Terraform) to enable safe, repeatable, fully automated infrastructure releases.
Drive fleet-wide initiatives such as cluster lifecycle automation, upgrade orchestration, policy enforcement, observability improvements, and disaster recovery readiness.
Partner with SRE, security, and product engineering teams to understand operational bottlenecks and deliver self-service platform capabilities that reduce toil and accelerate delivery.
Be an AI amplifier: embed AI tooling into everyday engineering and operations workflows - using agentic coding assistants (e.g., Claude Code) for development, code review, infrastructure automation, and intelligent operational runbooks - and help raise the team's AI fluency and force-multiply its output.
Bring an agentic mindset to platform and operational problems: identify toil and repetitive operational work, and design agent-driven or AI-augmented automations that let the platform operate, observe, and heal itself with minimal human intervention.
Participate in design reviews, write clear technical documentation, operational playbooks, and RFCs, and mentor junior engineers on platform, operations, and cloud-native best practices.
Contribute to on-call rotations and continuously improve the operational posture, monitoring, and incident response capability of the platform through automation and post-incident learning.
You're Our Person If...- 5+ years of professional experience in cloud infrastructure engineering, operations, and continuous deployment.
- Strong, hands-on Kubernetes experience - operating, troubleshooting, and automating clusters in high-scale production environments (controllers/operators, networking, scaling, upgrades).
- Good programming skills in Golang and/or Python, with experience building production-grade services, CLIs, or infrastructure automation tooling.
- Multi-cloud infrastructure and operations experience, primarily AWS, with working knowledge of core compute, networking, IAM, storage, and managed Kubernetes services.
- Hands-on Git Ops experience with Flux or Argo CD, and infrastructure-as-code experience with Pulumi (or comparable tooling such as Terraform).
- Demonstrated AIOps and automation fluency - you actively leverage AI tools and agentic workflows to build closed-loop automations, accelerate root-cause analysis, and design intelligent self-healing systems.
- An agentic operations mindset - you don't just use AI as a chat box; you know how to delegate complex engineering and operational tasks to AI agents, orchestrate multi-agent workflows for incident response, and use AI to drastically reduce MTTR (Mean Time to Resolution).
- Strong communication and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).