AI Engineer - Model Post Training
Listed on 2026-10-02
-
IT/Tech
Machine Learning/ ML Engineer, AI Evaluation, AI Engineer (Applied/Software)
Wider Security LLC is hiring a part-time
, fully remote AI Engineer to focus on post-training and alignment for large language models work centers on supervised fine-tuning and safety-focused evaluation, with hands‑on experience needed to support model tuning from 7B through 70B and beyond
.
This role is built for an async-first team and emphasizes measurable outcomes across training, calibration, and adversarial robustness. You will help lead end-to-end post-training workflows, from data preparation through evaluation and deployment-oriented integration.
What you’ll do- Lead and contribute to post-training workflows including supervised fine-tuning
, instruction tuning,
DPO
, RLHF
, RLAIF
, and related alignment techniques - Use QLoRA and other efficient fine-tuning methods across models spanning 7B to 70B+ parameter ranges
- Train models to produce reliable structured outputs under adversarial input conditions
- Build and run evaluation pipelines for safety-critical behavior, including adversarial test suites, red-team integration (for example,
Garak
), and regression tracking across model versions - Calibrate decision thresholds to support tiered policy configurations, including logprob-based confidence calibration at the serving layer
- Design training approaches that preserve inference-time policy specification
, enabling behavior changes without retraining - Curate and prepare training data
, evaluation sets
, and preference data pipelines - Iterate on training strategy to improve task performance, calibration, and adversarial robustness
- Document approaches and decisions clearly for a distributed, async-first team
- US citizenship and current residency in the United States (firm requirement)
- Concrete, verifiable production experience post-training open-weight LLMs, with experience at 7-8B
, 13-30B
, or 30B+ scales welcomed, and larger-scale experience a plus - Experience with modern open-weight model families such as Llama
, Qwen
, Mistral
, or similar - Hands‑on experience with QLoRA
, LoRA
, and efficient fine-tuning methods for large models - Background in alignment approaches including SFT
, DPO
, RLHF
, RLAIF
, or constitutional AI approaches - Experience training models for reliable structured output (JSON, schema-constrained generation, function-call style outputs)
- Familiarity with serving stacks such as vLLM
, TGI
, or similar, plus comfort working with logprob-level model outputs - Deep familiarity with distributed training frameworks including Deep Speed
, FSDP
, Megatron-LM
, or similar - Proficiency in Python and comfort with multi-GPU
, multi-node training infrastructure - Ability to work independently and manage time effectively in a part-time
, async-first environment
- Python
, QLoRA
, LoRA
, SFT
, DPO
, RLHF
, RLAIF - constitutional AI
, vLLM
, TGI
, Deep Speed
, FSDP
, Megatron-LM - Llama
, Qwen
, Mistral
, Garak
Applicants must be US citizens currently residing in the United States
. The company is unable to consider applicants outside the US or without US citizenship, regardless of work authorization status.
- Direct experience training safety classifiers, content moderation models, jailbreak or prompt injection detectors, or other trust-and-safety ML systems
- Experience with adversarial evaluation frameworks such as Garak
, promptfoo
, or similar - Comfort with deployment constraints typical of regulated or restricted-network environments
- Published research or open-source contributions related to LLM training, alignment, or AI safety
- Prior work at an AI lab, a foundation model team, or on a production safety classifier
- Experience designing or operating tiered policy systems where model behavior can be modulated at inference time
Location: Newport, RI (remote)
Job type: part time
Pay: USD 100 - 250 per hour
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).