Senior Software Engineer – Data & ML Platform
Listed on 2026-10-09
-
Software Development
DevOps, Cloud Engineer - Software, AI Engineer (Applied/Software), Machine Learning/ ML Engineer
We're looking for a Senior Software Engineer to take ownership of the production systems and internal data platform that power our Machine Learning (ML) and Operations Research (OR) work. This is a hands‑on, high‑ownership role at the intersection of backend engineering, cloud infrastructure, and data engineering. You'll work closely with our ML and OR specialists to ensure models and optimization solutions can move reliably from experimentation into production.
You'll inherit existing production systems and have the opportunity to improve and evolve them over time—from architecture and infrastructure to deployment, observability, and developer tooling. As our needs grow, you'll also help shape the roadmap for our internal data and ML platform. Because we're a small team, you'll have meaningful autonomy. You'll be the primary owner of these systems, collaborate directly with technical specialists and product teams, and have significant influence over architecture, tooling, and engineering practices.
You’ll Do Own and Evolve Production Systems
- Own, operate, and improve backend services running in Azure, including serverless services, batch workloads, and ML inference endpoints.
- Manage deployments and reliability across environments, including CI/CD, monitoring, alerting, incident response, and operational runbooks.
- Operate and optimize cloud compute environments, including autoscaling, container images, identity, and resource management.
- Improve the reliability, scalability, and maintainability of existing production systems over time.
- Design and build reliable ETL/ELT pipelines that transform data from relational and document databases into analysis‑ready datasets.
- Help develop our lakehouse‑style analytical layer and the infrastructure that supports it.
- Build and maintain infrastructure‑as‑code across environments using tools such as Terraform or Bicep.
- Manage cloud infrastructure including storage, application hosting, identity and access management, Key Vault, and cost optimization.
- Implement monitoring, logging, data‑quality checks, and freshness alerting across data workflows.
- Ensure data is handled securely through appropriate access controls, secrets management, and responsible treatment of sensitive information.
- Build new internal services, APIs, and developer tooling as the team's needs evolve.
- Build the infrastructure and tooling our ML and OR specialists need for experimentation, deployment, evaluation, and reproducibility.
- Turn research prototypes into reliable, production‑ready services and workflows.
- Own model packaging, versioning, deployment, and CI processes for ML and OR codebases.
- Build automated evaluation and benchmarking pipelines to monitor model performance, drift, and system reliability.
- Partner with Data Scientists and OR specialists to run and operationalize experiments.
- Establish strong practices around code quality, automated testing, version control, and CI/CD.
- Conduct peer code reviews and help teammates adopt scalable engineering practices.
- Improve existing systems incrementally rather than rebuilding for the sake of rebuilding.
- Help ensure our codebases remain maintainable and releasable as the team and platform grow.
- Work closely with Data Science, Operations Research, Product, and Engineering to integrate ML and optimization solutions into our products.
- Contribute to technical design discussions and decisions around architecture, scalability, reliability, and performance.
- Translate technical and business needs into pragmatic engineering solutions.
- Bachelor’s or Master’s degree in Computer Science, Software Engineering,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).