Senior Data Scientist (UAE
Abu Dhabi, UAE/Dubai
Listed on 2026-08-03
-
IT/Tech
Data Scientist, Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Data Engineering
CloudPSO is a Information Technology Outsourcing (ITO) company that assists in the acquisition of qualified staff to address complex digital problems in order to increase efficiency, reduce costs, and maintain compliance.
CloudPSO was founded in 2017 with an aim to provide businesses with a competent and skilled workforce at any given point in time and from any geographic region.
We are a US-based company with headquarters in Dallas (Texas) and a center of excellence in Pakistan. We have over 200 facility seats with an additional Work-From-Home facility. CloudPSO has skillful in-house software development teams with state-of-the-art tools, the latest VOIP technology platform, and secure infrastructure.
Our core values consist of client satisfaction, commitment, quality, and transparency.
We, at CloudPSO, hunt, analyze, recruit, train, and retain top-notch talent for you to help achieve your business goals. Optimizing mission-critical and day-to-day enterprise IT operations, CloudPSO enables businesses to transform, innovate and scale.
Job DescriptionThis is a remote position.
- Location: Remote - UAE
- Requirement: A Valid UAE work permit/employment visa is mandatory.
- Data Exploration and Analysis: Query and analyse large domain- or topic-specific data sets from both structured and unstructured sources, identify patterns and features. Ensure data meets quality standards and requirements before model development.
- Regulation Text Interpretation: Design and fine-tune Large Language Models (LLMs) to parse complex regulatory texts (e.g., building codes, military standards) and extract structured rules for automated compliance checking.
- Rule Formalization: Convert interpreted regulations into computer-processable formats (e.g., object-property-condition-value tuples) that can be executed by downstream compliance engines.
- Querying via NLP: Architect methods for LLMs to map natural language requirements directly to specific metadata entities within various schemas (e.g., mapping "systems design" to specified attributes).
- RAG Architecture: Implement Retrieval-Augmented Generation (RAG) pipelines that allow systems to query vast repositories of technical documentation and historical project data with high accuracy and low hallucination rates.
- Forecasting Engines: Develop time-series forecasting models to predict spend categories and material demand by correlating internal ERP data with external macroeconomic signals.
- Classification & Risk Scoring: Build machine learning classifiers to categorize supplier risks and operational anomalies, integrating data from diverse sources to create dynamic risk scores.
- Data Extraction Pipelines: Design robust pipelines to extract and transform raw data (from Data Lakehouse, external web sources, or SAP and other databases) into features required for predictive modeling and automated rule checking.
- Model Orchestration: Work with Back End Engineers to integrate AI models into a cohesive "compliance engine" or "risk engine" that can be invoked programmatically via robust APIs.
- Optimization: Streamline model performance to ensure complex checks (e.g., analyzing large datasets or processing thousands of supplier records) can be executed within reasonable time frames, potentially using batching or asynchronous processing.
- Quality Assurance: Validate model outputs against known test cases and historical data, debugging false positives/negatives to refine algorithms and ensure "defense-grade" reliability.
- Core AI/ML: Expert proficiency in Python and standard ML libraries (Tensor Flow/PyTorch, Scikit-learn, Pandas, Num Py). Strong grasp of both supervised and unsupervised learning techniques.
- NLP & LLMs: Deep experience with transformer-based models (GPT, BERT, Llama) and prompt engineering techniques (few-shot learning, fine-tuning) for domain-specific tasks.
- Data Engineering: Proficiency in handling complex data structures (JSON, XML) and familiarity with database querying (SQL/No
SQL) or graph data structures.
Experience with data extraction from specialized formats is a significant plus. - Backend Awareness: Understanding of how to expose models via RESTful APIs (Flask/FastAPI) and integrate them into larger software architectures.
- Statistics: Solid understanding of statistics, probability distribution, A/B testing. Adept at identifying and mitigating biases in datasets
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).