Senior AI Platform Engineer
Job in
Richardson, Dallas County, Texas, 75080, USA
Listed on 2026-08-02
Listing for:
Jobtailor
Full Time
position Listed on 2026-08-02
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
- Define standards-based deployment patterns for responsible AI agents, reusable platform capabilities, and secure agent runtime architectures, including identity, access, and secrets management.
- Automate AI platform buildout, release standardization, environment provisioning, CI/CD, Terraform/Helm deployments, and operational runbooks.
- Operate reliable AI platform services with incident response, capacity planning, Open Telemetry traces/logs/metrics, monitoring, rollback, and disaster recovery practices.
- Collaborate across Info Sec, Cloud Ops, Dev Ops, and platform engineering teams to create responsible agentic factory standards for guardrails, governance, secure releases, observability, and production readiness.
- Coach engineers, lead design reviews, guide implementation toward approved architecture patterns, and drive practical tradeoffs across cost, speed, reliability, and security.
- 8 to 10 years of engineering experience, including 7+ years building and operating production systems on cloud platforms, plus hands-on AI/ML service deployment in production.
- 2-3 years using AI coding tools (Claude Code, Codex, Cursor) to automate platform buildout, deployment, testing, troubleshooting, and documentation.
- Strong experience with GCP/AWS or hybrid cloud/datacenter deployments;
Docker, Kubernetes, Git Hub Actions with reusable workflows and self-hosted runners, Terraform, and Helm. - Working knowledge of authentication and authorization (OAuth 2.0, OIDC, SAML, JWT, RBAC, IAM), including workload identity, service-to-service auth, and securing API and tool access.
- 4+ years scripting and automation experience, preferably Python and JavaScript, with strong troubleshooting across Linux, containers, Kubernetes, networking, and production incidents.
- Ability to architect secure, cost-efficient hosting for open-source LLMs, on-prem or in dedicated cloud, as an alternative to commercial model APIs.
Demonstrates expertise in defining deployment patterns for responsible AI agents and automating AI platform buildout, with a strong focus on secure architectures and incident response. Proven ability to collaborate across teams and coach engineers in implementing best practices for production readiness.
Highest-signal resume keywords- AI/ML Service Deployment
- Cloud Platform Engineering
- Terraform and Helm
- Authentication and Authorization
- Scripting and Automation
- AI Coding Tools
- Production Systems Engineering
- Incident Response
- Capacity Planning
- Monitoring and Observability
- Scripting in Python
- Scripting in Java Script
- Troubleshooting
- Deployment Automation
- Secure Architecture Design
- Coaching
- Collaboration
- Leadership
- Guidance
- Responsible AI
- Production Readiness
- Governance
- Guardrails
- Disaster Recovery
- GCP
- AWS
- Docker
- Kubernetes
- Git Hub Actions
- Open Telemetry
- CI/CD
- Linux
- RBAC
- IAM
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×