Systems Development Engineer
Listed on 2026-09-03
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
About Us
At Union, we are solving one of the hardest challenges in AI infrastructure today: enabling high-velocity iteration while maintaining seamless production-readiness for AI workloads at scale.
Flyte, the open-source project we steward, has emerged as the modern standard for data and AI orchestration, and is trusted by leading technology organizations including Linked In, Stripe, and Wayve to run millions of mission-critical workflows on the platform. These workflows comprise data preparation, model training, and scaled inference spanning thousands of GPUs, all major clouds, and on-premise infrastructure.
We have a technical founding team who created Flyte while at Lyft, a deep bench of infrastructure experts from top companies, and have raised from top investors like NEA and Nava Ventures.
About the RoleWe are hiring a Systems Development Engineer to improve the reliability, operability, and customer experience of our production platform. This is a 50/50 operations and engineering role. Part of your time will be spent investigating customer-impacting production issues, and part will be spent building the tools, automation, system design, and engineering practices that prevent those issues from recurring.
For this role, production is the customer. You will work from real operational signals: customer issues, incidents, on-call pages, recurring support patterns, and gaps in observability or automation. You will sit in engineering and partner with customer-facing teams to turn those signals into durable platform improvements.
This is not a traditional support role. It is a systems engineering role for someone who can debug deeply, communicate clearly, and write software that reduces operational load. The full engineering team backs you on on-call.
This role is hybrid, based out of our Seattle office.
What You'll DoInvestigate and resolve customer-impacting production issues across cloud infrastructure, workflow execution, access control, storage, networking, deployment systems, and observability.
Identify patterns in customer issues and convert them into automation, product improvements, runbooks, tests, or design changes.
Build internal tools and diagnostics that make production issues easier to detect, understand, and resolve.
Improve platform observability, including logs, metrics, dashboards, alerts, and customer-visible debugging information.
Participate in design and development so systems are easier to operate, debug, and support, and guide engineering teams toward durable fixes.
Define and uphold operational engineering practices: production readiness, alert quality, runbook discipline, observability standards, regression prevention, and code quality.
Drive measurable reductions in on-call pages, recurring customer issues, manual operational work, and time-to-resolution.
Strong software engineering skills in Python, Go, Java, or a similar language.
Experience debugging production systems across multiple layers of the stack.
Practical knowledge of Kubernetes, Linux, cloud infrastructure, distributed systems, networking, storage, and IAM.
Experience with infrastructure-as-code, deployment systems, CI/CD, observability, and operational automation.
Ability to move from ambiguous customer symptoms to clear technical diagnosis and durable remediation.
Strong judgment about when to fix directly, automate, elevate, redesign, or build a broader platform improvement.
Clear written and verbal communication, especially around root cause analysis, technical recommendations, runbooks, and design feedback.
A bias toward reducing toil through engineering rather than repeatedly solving the same issue by hand.
Operating customer-facing SaaS, cloud infrastructure, self-hosted or on-prem deployments, or workflow orchestration systems.
Batch workloads, autoscaling, capacity management, identity and access systems, storage systems, or platform observability.
Improving on-call health, reducing ticket volume, or building production diagnostics.
Working across support, customer success, product, and engineering teams.
Customer issues are diagnosed faster and recur…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).