×
Register Here to Apply for Jobs or Post Jobs. X

Senior AI Engineer, Foundation Model Training, SeekrGEO

Job in Austin, Travis County, Texas, 78701, USA
Listing for: Seekr
Part Time, Apprenticeship/Internship position
Listed on 2026-07-20
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Job Description & How to Apply Below

AI Engineer, Foundation Model Training, SeekrGEO

Austin, Texas, United States;
Reston, Virginia, United States

Seekr's Mission

Seekr builds trusted AI for mission-critical decisions. Our platform helps organizations build, govern, and deploy secure, explainable AI rooted in their own data across cloud, on-premises, edge, and air-gapped environments. We care deeply about transparency, auditability, and defensibility because high-stakes AI is only useful when people can understand and trust how it behaves.

About the Opportunity

SeekrGEO is Seekr's geospatial AI product. This role contributes to the foundation model program behind it: pretraining and post-training of large multi-modal models on geospatial data, together with the distributed training systems that make that work possible  focus is training, but the role supports the full model lifecycle through deployment.

As a Research Engineer you lead the training systems that make ambitious model programs possible: large-scale distributed training, parallelism strategies, data infrastructure, and the operational rigor that multi-week runs demand. You will work alongside Research Scientists on modeling and recipe decisions, and you translate ideas from research papers into working code and decide whether they deserve a full training run.

What You'll Do
  • Build and harden training infrastructure on accelerator clusters: data loaders, parallelism strategies, checkpointing, fault tolerance, and the evaluation harness that catches regressions before customers do.
  • Own the parallelism strategy for our training workloads: FSDP, tensor / pipeline / sequence parallelism, ZeRO variants, activation and gradient checkpointing, mixed precision, and the memory and throughput tradeoffs that come with each.
  • Diagnose distributed training failures and turn fixes into reusable platform improvements.
  • Design and operate the data pipeline for large training corpora: sharded formats, streaming loaders, deduplication, mixture tuning, and the versioning discipline that makes runs reproducible.
  • Keep multi-week training runs healthy through checkpoint management, fault-tolerant and elastic training, and the operational hygiene needed for long-horizon runs on shared infrastructure.
  • Do performance work on accelerators: kernel-level profiling, attention kernel selection and tuning, memory layout optimization, and closing the gap between theoretical and observed throughput.
  • Build the evaluation infrastructure that makes model comparisons trustworthy and reproducible, both during training and after deployment.
  • Support deployed models through their lifecycle: monitor systems behavior in production, diagnose regressions, and close the loop back into the next training cycle.
  • Contribute improvements back to Seekr Flow training so the platform gets stronger with every run.
  • Partner with Research Scientists to pressure-test ideas: reproduce a paper's core claim, verify a proposed recipe scales, and turn research prototypes into production runs.
  • Partner with the SeekrGEO product team and customer-facing teams to align training infrastructure with the workflows the model needs to support.
  • Use AI coding assistants effectively as part of a modern engineering workflow while maintaining strong judgment over training code, systems code, and infrastructure.
What We're Looking For
  • Strong background in modern ML systems, with deep familiarity with transformer architectures, multi-modal models, and the practical realities of training them at scale.
  • Fluency with PyTorch and the distributed training ecosystem (FSDP, tensor / pipeline / sequence parallelism, ZeRO, checkpointing strategies).
  • Hands-on experience with at least one large-scale training framework such as Megatron-LM, torch titan, or Deep Speed.
  • Ability to move comfortably between engineering and research. You can read a paper, reproduce its core idea, and pressure-test whether it will hold up at scale.
  • Demonstrated experience contributing to a large model run through pretraining or continued pretraining, not just fine-tuning a frontier checkpoint.
  • Comfort designing experiments and evaluating ambiguous technical tradeoffs.
  • Strong Python and software engineering fundamentals, with comfort in testing, code review, CI/CD, debugging, and performance analysis.
  • Fluency with AI coding assistants and the modern developer workflows they enable.
  • Clear communication and strong collaboration across technical and non-technical partners.
  • Reside near Austin, TX or Reston, VA and able to work 3 days per week in office.
Preferred Qualifications
  • Experience operating distributed training at scale across accelerator clusters, with comfort in collective communication and the failure modes specific to large-scale runs.
  • Hands-on experience with Megatron-LM, torch titan, and other distributed training frameworks.
  • Performance work on accelerators: kernel-level profiling, mixed precision, activation and gradient checkpointing, attention kernels, memory layout optimization.
  • Experience with AMD ROCm is a strong…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary