More jobs:
Technical Architect
Job in
Toronto, Ontario, C6A, Canada
Listed on 2026-09-03
Listing for:
Mphasis
Full Time
position Listed on 2026-09-03
Job specializations:
-
IT/Tech
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Cloud Computing: Infrastructure & Operations, Data Engineering
Job Description & How to Apply Below
MLOps Engineer (Cloud Engineer + Dev Ops)
Location:
Canada Job Overview
We are looking for an experienced
POSITION /TITLE:
MLOps Engineer (Cloud Engineer + Dev Ops)
Location:
Canada Job Overview
We are looking for an experienced MLOPs / LLMOPs Engineer with a strong background in deploying and monitoring machine learning and large language model (LLM) pipelines. The ideal candidate will have 10+ years of experience in MLOPs, with expertise in setting up end-to-end ML/LLM pipelines using open-source tools and cloud-native solutions on platforms like AWS, GCP, and Azure
. This role requires hands-on knowledge in deploying, automating, and monitoring ML/LLM workflows, with a solid grounding in Dev Ops practices to ensure seamless CI/CD processes.
- Pipeline Design & Implementation:
- Design, build, and manage MLOPs and LLMOPs pipelines for data ingestion, model training, validation, deployment, and monitoring.
- Use open-source tools such as Mlflow, Kubeflow, DVC, and Airflow to automate and monitor machine learning workflows.
- Implement scalable LLM-specific solutions for model training and inference, optimizing resource allocation and deployment efficiency.
- Cloud-native MLOPs Implementation:
- Set up and manage MLOPs pipelines in Primary in GCP (Sage Maker, EKS, Lambda, S3), or have similar experience with GCP (Vertex AI, AI Platform Pipelines), and Azure (Machine Learning, AKS, Azure Functions).
- Manage model versioning, retraining, and deployment workflows on cloud platforms to ensure consistent performance and availability.
- Execute CI/CD pipelines for ML models with Git Hub Actions, Jenkins, or Git Lab CI.
- Model Monitoring & Performance Optimization:
- Monitor models in production using Prometheus, Grafana, and Tensorboard, establishing observability metrics for model drift, accuracy, and latency.
- Collaborate with Data Engineering and ML teams to implement scalable and efficient pipelines
- Use A/B testing and shadow deployment strategies to validate and optimize LLM model performance in real-time.
- LLM-specific Model Operations:
- Deploy and monitor LLMs for specific tasks, ensuring they adhere to performance SLAs and are optimized for cost.
- Understand techniques of fine-tuning, optimizing inference, and managing infrastructure costs for large LLMs.
- Proficiency with Kubernetes and Docker for container orchestration and model deployment.
- Experience with open-source MLOPs tools (Mlflow, Kubeflow, DVC) and data versioning.
- Hands-on experience with cloud-native ML tools in AWS, GCP, or Azure and associated ML services.
- Knowledge of Python or Bash scripting for automating processes and custom integrations.
- Solid understanding of CI/CD practices and tools like Git Hub Actions, Jenkins, or Git Lab CI/CD to build and deploy ML/LLM models.
- Proficient in infrastructure-as-code tools, such as Terraform or Ansible, to enable automated provisioning and configuration management.
- Python
- SQL, No-SQL, PySpark (Optional)
- Supervised, Unsupervised Learning & Model evaluation metrics
- NLP, RAG, GenAI, LLMs
- Deep Learning (Sequential & Functional APIs) using Pytorch/Tensor Flow
- MLOPs & Mlflow Experiment Tracking
- Explainable AI (XAI) LIME SHAP(Optional)
- (Primary: AWS) or Handons Expertise on any of cloud platforms
- Azure AI/ML,
- Google Vertex AI,
- Databricks Studio
- Bachelors/ Master’s/PhD Degree in Mathematics, Statistics, Physics, Computer Science, Engineering, Data Science, or a related relevant degree from quantitative field.
- Understanding of Agile and Scrum methodologies.
- Ability to follow SDLC processes and contribute to technical documentation.
- Must have Structural thinking and goal-oriented approach to problem-solving
- Self-motivated and capable of working independently with minimal management supervision.
- Well-developed design, analytical & problem-solving skills
- Excellent communication and interpersonal skills.
- Excellent team player, able to work with virtual teams in several time zones.
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×