Data Scientist; Entry Level to SME) TS/SCI Poly Security Clearance
Job in
Arlington, Arlington County, Virginia, 22201, USA
Listed on 2026-08-28
Listing for:
CGI
Full Time
position Listed on 2026-08-28
Job specializations:
-
IT/Tech
Machine Learning/ ML Engineer, Data Scientist, Data Analyst, Data Engineering
Job Description & How to Apply Below
Position
Description:
CGI Federal has an exciting opportunity for Data Scientists within our Intel sector advancing the national security mission through cutting edge technology. You must have a passion for keeping pace with rapidly evolving technology advancements and leveraging your knowledge on a highly collaborative team to deliver state-of-the-art capabilities. The Data Scientist supports the analytics team by cleaning messy data, running queries, and helping to identify business trends.
They focus on foundational tasks like Exploratory Data Analysis (EDA) and assist senior staff in preparing datasets for predictive models. CGI Federal is growing its high-performance team whose members share a passion for building high-quality, scalable, advanced IT solutions in a collaborative, fast-paced, outcome-driven mission. This position is located in our Arlington office; however, a hybrid working model is acceptable.
Your future duties and responsibilities:
Key Responsibilities (Entry Level)
• Data Preparation:
Use SQL to extract raw data and Python or R to clean and format datasets.
• Exploratory Analysis:
Analyze data to find trends, patterns, and anomalies.
• Visualization:
Create basic charts and dashboards using tools like Tableau or Python libraries to present findings to the team.
• Model Assistance:
Help senior data scientists train, test, and evaluate basic machine learning algorithms.
Key Responsibilities (Junior Level)
• Data Wrangling:
Clean, process, and validate raw structured and unstructured data to ensure uniformity and accuracy.
• Exploratory Data Analysis (EDA):
Analyze data to identify trends, patterns, and anomalies.
• Modeling:
Assist in developing, testing, and updating basic machine learning models and statistical algorithms.
• Visualization & Communication:
Build dashboards and presentations to clearly communicate findings and recommendations to both technical and non-technical stakeholders.
• Pipeline Maintenance:
Collaborate with data engineers and senior data scientists to maintain and optimize data pipelines
Key Responsibilities (Mid-Level)
• Model Development & Deployment:
Design, train, evaluate, and deploy robust machine learning and deep learning models to solve ambiguous business problems.
• Advanced Analytics:
Conduct rigorous exploratory data analysis (EDA) and apply complex statistical techniques (e.g., A/B testing, regression analysis, clustering) to extract deep insights.
• Data Engineering & Pipelines:
Extract, clean, and manipulate unstructured datasets across distributed systems. Contribute to the design and optimization of data pipelines.
• Stakeholder
Collaboration:
Translate high-level business goals into precise data science requirements. Present actionable recommendations to both technical and non-technical stakeholders.
• Technical Leadership:
Act as a subject matter expert and mentor junior analysts or entry-level data scientists on methodology and coding best practices.
Key Responsibilities (Senior Level)
• Advanced Modeling:
Architect and deploy Deep Learning (DL), Natural Language Processing (NLP), and Large Language Models (LLMs).
• System Scalability:
Build and optimize distributed data processing pipelines (e.g., using Spark) and automate reproducible workflows.
• Translational Strategy:
Convert ambiguous, high-dimensional business/mission requirements into strict technical requirements.
• Leadership & Mentorship:
Lead end-to-end projects autonomously and mentor junior data scientists and engineers
Key Responsibilities (SME Level)
• Advanced AI & Model Architecture:
Architect and operationalize complex Agentic AI systems and customized Retrieval-Augmented Generation (RAG) frameworks to support dynamic domain requirements.
• Constrained Environment Optimization:
Optimize resource-heavy machine learning models and deep neural networks to perform reliably at the edge or within restricted computing ecosystems.
• Unstructured Data Synthesis:
Develop reproducible analytical models and quantitative techniques to draw definitive conclusions from incomplete, noisy, or highly unstructured raw datasets.
• Pipeline Automation:
Design and automate scalable data pipelines,…
Position Requirements
Less than 1 Year
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×