Data Engineer, AWS Glue and Data Lake (GovCloud
Listed on 2026-08-16
-
IT/Tech
AWS, Data Engineering
Data Engineer, AWS Data Lake (Gov Cloud, IL5)
DATA & ANALYTICS | ACTIVE SECRET CLEARANCE REQUIRED | ACTIVE CAC REQUIRED | U.S. CITIZENSHIP REQUIRED | FULLY REMOTE (CONUS)
ABOUT THE ROLE
the company is sourcing a cleared Data Engineer to design and operate a multi-account data lake inside AWS Gov Cloud (US), at Impact Level 5. The work sits where storage, movement, and governance meet. You will land raw data, shape it into something engineers and analysts can actually query, and keep the pipelines that do it running clean under a controlled unclassified boundary.
Ingestion is only half the job. The other half is making the environment repeatable, so that infrastructure ships as code, account baselines are hardened by pipeline rather than by hand, and compliance evidence is produced as a byproduct of delivery instead of reconstructed after the fact.
This seat carries an engineering simulation workload, which changes the shape of the pipeline. Finite element analysis datasets arrive large, arrive often, and punish a lake that was tuned for tidy tabular records. If you are comfortable in a Terraform module, comfortable debugging a Glue job that died at three in the morning, and equally comfortable explaining a partitioning decision to a data consumer who does not care how it was built, this posting was written for you.
WHATYOU WILL DO
Architect the lake. Design and implement multi-account AWS data lake architectures within Gov Cloud (US), including S3 storage layers, partitioning strategy, and catalog structure that hold query cost and query time steady as volume grows.
Control access by classification. Configure data access controls that enforce classification level and need to know, and apply data classification tagging and encryption at rest to IL5 requirements.
Own the ETL. Develop, schedule, and tune AWS Glue jobs and crawlers for both structured and unstructured ingestion, then carry the surrounding workflow end to end, from source through curated output.
Feed the simulation workload. Configure data asset pipelines supporting finite element analysis inputs and outputs, and optimize high-throughput processing for large simulation file I/O such as FEMAP and Simcenter Nastran datasets.
Ship infrastructure as code. Author and maintain Terraform modules for baseline account infrastructure and networking, and configure account baseline hardening to IL5 security requirements through code rather than console.
Automate the deployment path. Build deployment pipelines for infrastructure and application code, including Terraform state management in restricted or air-gapped environments where the usual backend patterns do not apply.
Guard data quality. Implement data quality monitoring against operational standards, build validation and alerting into the pipelines, then act on what they surface.
Document what you build. Maintain data flow diagrams, runbooks, and schema documentation that survive your absence.
Partner across the mission. Work with engineering, analytic, and security stakeholders to translate data requirements into working pipelines, and coordinate directly with government technical stakeholders.
REQUIRED QUALIFICATIONSActive Secret clearance, and personnel eligibility appropriate to IL5 Controlled Unclassified Information work.
Active Common Access Card, required for environment access and for coordination with government stakeholders. Candidates without a current CAC should expect the sponsorship timeline to be the critical path on this task.
U.S. citizenship.
Demonstrated experience operating within AWS Gov Cloud (US) at Impact Level 5.
Hands-on experience designing, building, or operating a data lake in AWS, including S3 storage design, partitioning, and cataloging.
Proficiency with multi-account AWS architecture patterns and baseline security controls.
Demonstrated ETL and data processing depth with AWS Glue, including job development, crawlers, and the AWS Glue Data Catalog.
Infrastructure as code experience, with Terraform strongly preferred, covering module authorship, provisioning, and state management.
Python or PySpark for data transformation work.
Comfort operating independently in a fully…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).