Platform Engineer IV
Listed on 2026-07-31
-
IT/Tech
AWS, Cloud Computing: Infrastructure & Operations, Data Engineering, SRE/Site Reliability
Employment Type: Long-term W2 contract with renewals typically occurring every 3–6 months. Contract extensions and conversion to a full-time position are possible but are not guaranteed.
Pay Range: $70–$75/hour, depending on experience
Start Date: ASAP
Location: Remote eligible during the contract term. Preference will be given to candidates who are open to working onsite 1–2 days per week in one of the following locations:
- St. Louis, Missouri
- Charlotte, North Carolina
- Stamford, Connecticut
If the position converts to full-time employment, the expected schedule would be four days onsite and one day remote.
Position Overview
The Infrastructure Intelligence and Analytics team builds and operates the data platform and AI-agent infrastructure that supports proactive network monitoring and autonomous investigation across large-scale network operations.
The Platform Engineer IV will design, build, automate, and maintain the AWS infrastructure supporting the organization’s data lake, data pipelines, AI-agent runtime environments, CI/CD processes, graph database systems, and production applications.
This position combines cloud infrastructure engineering, Dev Ops, automation, application development, and production support. The selected candidate will help ensure environments are stable, scalable, secure, and capable of supporting data science and agentic AI workloads at enterprise scale.
Responsibilities may span several focus areas depending on team priorities and the selected candidate’s strengths. Candidates are not expected to have experience with every technology listed below.
Key Responsibilities
AWS Infrastructure and Data Platforms
- Design, build, and manage AWS infrastructure supporting an enterprise data lake, including Amazon S3, AWS Glue, Athena, and EMR.
- Manage secure cross-account connectivity and data movement between internal teams, upstream data providers, and downstream consumers.
- Configure and support VPC networking, security groups, IAM roles and policies, Private Link, VPC peering, and transit gateway connectivity.
- Build and maintain infrastructure supporting AI-agent runtime environments, including compute resources for agents deployed through enterprise AI platforms.
- Design event-driven infrastructure using triggers, queues, messaging platforms, and publish/subscribe frameworks.
- Support the deployment and operation of AWS Neptune for network-topology and digital-twin use cases.
- Implement and maintain infrastructure as code using Terraform or Cloud Formation.
- Manage secrets, access-key rotations, service accounts, and security controls using tools such as AWS Secrets Manager and Delinea.
- Support secure cloud and on-premises integrations, including Splunk and related platforms.
CI/CD and Deployment Automation
- Build and maintain CI/CD pipelines using Git Lab CI/CD, Docker, Artifactory, and related deployment technologies.
- Manage container builds, image versioning, deployment promotion, and rollback processes.
- Automate the movement of applications and AI agents from development through production.
- Support cross-account deployments, AI-gateway integrations, and connectivity with internal platform teams.
- Improve deployment reliability, repeatability, auditability, and release speed.
Application Development and Internal Tooling
- Develop and maintain Python-based scripts, command-line tools, automation utilities, small services, and internal applications.
- Build integration code connecting internal systems with upstream data providers, downstream consumers, and external platforms.
- Support and enhance existing ETL pipelines written in Scala and Apache Spark.
- Assist with troubleshooting data pipelines and onboarding new data sources.
- Write maintainable, reusable, and well-documented code using established software-engineering practices.
- Develop unit and integration tests and incorporate automated testing into CI/CD pipelines.
Production Operations
- Maintain production stability through monitoring, alerting, troubleshooting, and incident response.
- Support service-level expectations for data-pipeline availability and AI-agent uptime.
- Implement health checks and monitoring for application errors, latency, resource utilization,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).