Job Description & How to Apply Below
Purpose of the Position: The AI Platform Engineer- MLOps/LLMOps will be responsible for designing, implementing, and operationalizing enterprise-grade AI, Machine Learning, and Generative AI solutions. The role will focus on deploying and monitoring AI applications, establishing governance frameworks, and enabling secure, reliable, and cost-effective AI operations across cloud environments. The individual will provide technical leadership and drive best practices for AI engineering, MLOps, and LLMOps initiatives.
Key Result Areas and
Activities:
Reliability and Availability of AI services and solutions
Design scalable AI/ML and Generative AI deployment architectures.
Define enterprise standards for model deployment, monitoring, and lifecycle management.
Design and operationalize RAG, agentic AI, prompt management, vector databases, and LLM integration frameworks.
MLOps & LLMOps Implementation
Establish CI/CD pipelines for machine learning and LLM-based applications.
Automate model training, deployment, versioning, and rollback processes.
Optimize LLM performance, scalability, and cost efficiency including Inference cost management
Model Governance & Reliability
Implement frameworks for model monitoring, evaluation, drift detection, observability, and compliance.
Ensure responsible AI and governance standards are followed.
Technical Leadership & Stakeholder Collaboration
Mentor AI engineers, junior platform engineers and data scientists.
Collaborate with business and technology teams to translate AI use cases into production-ready solutions.
Essential
Skills:
LLM serving and inference optimization and running an inference gateway that routes across providers and open-weight models with failover
LLM observability and tracing: end-to-end request tracing through retrieval, prompt, tool calls and generation using Lang Smith or Azure AI Foundry
Evaluation as a platform capability: building the harness product teams use to regression-test AI behaviour, run LLM-as-judge and offline evals, and gate releases on measured quality rather than judgement calls
Prompt, model and config lifecycle management: versioning prompts and system messages like code, safe rollout and rollback, A/B and shadow testing of model or prompt changes in production
Token and inference cost engineering: spend attribution by team, feature and tenant; caching and model-routing strategies; GPU utilisation and right-sizing
Retrieval infrastructure operations: running and tuning vector or hybrid search at scale, embedding pipelines, index refresh strategies, and retrieval quality monitoring
Kubernetes and GPU infrastructure in production:
Docker, autoscaling, node pools, GPU scheduling and sharing, plus infrastructure as code (Terraform or Bicep) and Git Ops (Argo CD or Flux)
Strong Python and the software engineering discipline to build self-service platform tooling and internal SDKs that engineers actually adopt
Classical MLOps foundations: model registry and versioning (MLflow, Azure ML, or Sage Maker), automated retraining and redeployment pipelines, drift and data quality monitoring, and pipeline orchestration with strong SQL
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×