Data Scientist - NLP
Listed on 2026-08-08
-
IT/Tech
Data Scientist, Machine Learning/ ML Engineer, Data Analyst
Analytica is seeking a Data Scientist to support long term federal client engagements projects in the DC Metro area. The role will apply statistical programming, modeling, visualization techniques, data mining, and forecasting skills to analyze challenging public sector problems.
This position is fully remote.
Analytica has been recognized by Inc for three consecutive years as one of the 250 fastest growing businesses. We offer competitive compensation with opportunities for bonuses, employer paid health care, training and development funds, and 401(k) match.
Responsibilities- Pre-processing:
Demonstrate skills and experience to collect, clean, and prepare data sets for input into a computational model using Python. Explain various methods applied such as stop word removal, stemming, lemmatization, and tokenization. - Feature engineering and attribute evaluation:
Demonstrate experience with NLP feature engineering methods such as TF-IDF, word2vec, Glo Ve, and Fast Text, identifying key determinants for modeling that exist in the business process and within existing data sets as well as selecting evaluation protocols. - Modeling:
Demonstrate experience selecting classification modeling techniques to fit the business problem. Examples include supervised and unsupervised learning, regression, neural networks and deep learning, natural language processing, etc. - Validation:
Describe experience with investigating, reporting, and justifying model results. - Visualization:
Experience presenting results of modeling activities, depicting insights realized, and explaining relevance to the organization’s business challenges.
- Master’s degree required and PhD preferred in Statistics, Mathematics, Computer Science, or similar.
- High degree of experience utilizing SAS, R, or Python to support NLP use cases such as Document Summarization, Named Entity Recognition, Sentiment Analysis, and/or Topic Modeling.
- At least four years of experience developing scalable, production-ready NLP solutions using scikit-learn, Keras, Tensor Flow, PyTorch, Spark NLP.
- Experience using git or Git Hub to version control source code.
- Experience leveraging transformer architecture to develop NLP models.
- Experience with open source NLP packages such as Gensim, Spa Cy, or NLTK.
- Experience with BERT, GPT-J, RoBERTa, T5 or other transformers.
- Experience with GenAI and Prompt Engineering is a plus.
- Experience in Databricks and MLflow is a plus.
- Experience with machine translation and transcription of foreign language documents using Microsoft Azure translation services is a plus.
- Experience working in an AWS cloud environment and with related AWS services such as Bedrock and Textract.
- Experience coordinating and maintaining user stories.
- Must be a U.S. citizen.
- Must be able to obtain and maintain a Public trust security clearance.
Analytica LLC is an Equal Opportunity Employer. We are committed to providing equal employment opportunities to all individuals, regardless of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, or any other characteristic protected by applicable federal, state, or local law. As a federal contractor, we comply with the Vietnam Era Veterans' Readjustment Assistance Act (VEVRAA) and take affirmative action to employ and advance in employment qualified protected veterans.
We ensure that all employment decisions are based on merit, qualifications, and business needs. We prohibit discrimination and harassment of any kind. Analytica LLC also provides reasonable accommodations to applicants and employees with disabilities, in accordance with applicable law.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).