Research Engineer II - ML Ops
Listed on 2026-07-28
-
Software Development
Job Description Summary
Are you passionate about building reliable cloud infrastructure that helps researchers innovate faster? In this role, you will work at the intersection of cloud engineering, machine learning operations, and research computing, helping teams develop, test, and deploy cutting‑edge solutions that advance MIM Research initiatives.
Are you passionate about building reliable cloud infrastructure that helps researchers innovate faster? In this role, you will work at the intersection of cloud engineering, machine learning operations, and research computing, helping teams develop, test, and deploy cutting‑edge solutions that advance MIM Research initiatives. You'll collaborate with researchers, engineers, and infrastructure partners to create scalable AWS‑based environments, develop internal tools, and build engineering solutions that enable impactful research.
We are looking for someone who enjoys solving complex problems, learning new technologies, and working in a collaborative, mission‑driven environment.
Key Responsibilities In This Role, You Will
- Partner with Dev Ops and infrastructure teams to migrate, optimize, and support research workloads on AWS cloud platforms.
- Design, build, and maintain machine learning operations (MLOps) pipelines that support model training, evaluation, deployment, and monitoring.
- Develop prototypes and internal tools that accelerate experimentation, model development, and research workflows.
- Translate research objectives into scalable, maintainable, and well‑documented engineering solutions.
- Promote and support engineering best practices, including:
- Code quality, testing, and reliability
- Documentation and version control
- Data management and governance
- Experiment tracking and reproducibility
- Effectively manage multiple projects while balancing fast‑paced research needs with long‑term engineering sustainability.
- Provide technical guidance and mentorship to early‑career engineers and support knowledge sharing across teams.
- Collaborate closely with research scientists, product teams, and infrastructure partners to deliver impactful solutions.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 2 to 4 years of experience in Dev Ops, Site Reliability Engineering (SRE), cloud infrastructure, or related technical roles.
- Experience supporting production systems within Amazon Web Services (AWS).
- Experience working with AWS services such as:
- Compute: EC2, ECS, EKS, Lambda
- Storage: S3, EFS
- Data:
DynamoDB - Machine Learning:
Sage Maker (training, pipelines, and deployment) - Experience with Infrastructure as Code (IaC) tools such as Ansible, Terraform, Cloud Formation, or AWS CDK.
- Familiarity with containerization and orchestration technologies such as Docker, Docker Compose, and Kubernetes.
- Experience building and maintaining continuous integration and continuous deployment (CI/CD) systems, including tools such as Git Hub Actions, Git Lab CI, or Jenkins.
- Strong foundation in Linux systems administration.
- Experience with monitoring and observability practices using tools such as Prometheus, Datadog, or similar technologies.
- Understanding of software engineering best practices, including:
- Software design patterns
- API development (REST and gRPC)
- Testing methodologies and maintainable code architecture
- Experience supporting research environments or collaborating closely with research teams.
- Ability to work effectively with evolving requirements, experimentation, and iterative development processes.
- Ability to lead technical initiatives involving multiple stakeholders and cross‑functional teams.
- Experience mentoring engineers and supporting the adoption of engineering best practices.
- Strong communication skills with the ability to connect technical concepts across research and engineering audiences.
- Experience with large‑scale distributed computing frameworks such as Spark or Ray.
- Background in high‑performance…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).