Senior Platform Engineer
Job in
Cambridge, Middlesex County, Massachusetts, 02140, USA
Listed on 2026-07-23
Listing for:
ICONSTAFF
Full Time
position Listed on 2026-07-23
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Job Description & How to Apply Below
Senior Platform/Infrastructure Engineer
Location: Fully remote (HQ Cambridge, MA)
Hours: 9–5 EST, with 2-day on-site visits every 6 weeks
You’ll be responsible for designing, scaling, and maintaining the infrastructure and internal developer platforms that power a real-time learning AI at a seed‑stage startup. The role blends infrastructure ownership with platform engineering to enable AI/product teams to ship quickly and reliably.
Key Responsibilities- Infrastructure
- Maintain production health: performance, reliability, cost efficiency, and security.
- Manage GCP Kubernetes clusters (GKE), networking, storage, and compute resources.
- Handle scaling, resource allocation, and high availability for growing customer demand.
- Refine observability: logs, traces, metrics, dashboards, and alerts.
- Perform security hardening and cost optimization.
- Platform Engineering
- Build internal tooling and abstractions for developer productivity.
- Design CI/CD pipelines using Git Hub Workflows and ArgoCD.
- Provide self‑service environments, internal portals, and deployment systems.
- Work closely with AI and full‑stack teams to optimize system architecture.
- Explain technical concepts and trade‑offs clearly to engineers and non‑engineers.
- Troubleshoot issues across multiple systems (Python, JavaScript, SQL).
- 5 years in production cloud environments at scale.
- Strong familiarity with GCP (primary) and some AWS experience.
- Experience with Kubernetes (GKE), node pools, and memory‑intensive jobs.
- Working knowledge of CI/CD systems (Git Hub Workflows, ArgoCD).
- Exposure to observability tools (Datadog), databases (Cloud SQL, Click House, Bigtable), and cloud services.
- Strong analytical and problem‑solving ability.
- Clear, collaborative communication.
- Curiosity and ownership mentality.
- Fluent in reading/debugging code across Python, JavaScript, SQL.
- Cloud: GCP (primary), AWS (secondary)
- Kubernetes: GKE, multiple node pools
- CI/CD: Git Hub Workflows, ArgoCD
- Data: Cloud SQL, Click House, Bigtable, GCS, Dataflow
- Networking: Cloudflare Workers, Durable Objects, Web Socket communication
- Monitoring: Datadog
- Environments: Production, Staging, Integration, Development
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×