×
Register Here to Apply for Jobs or Post Jobs. X

Machine Learning Engineer, Reliability

Job in Jacksonville, Duval County, Florida, 32290, USA
Listing for: DataJobs
Full Time position
Listed on 2026-10-02
Job specializations:
  • Software Development
    DevOps, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below

fal builds generative media model APIs, and this role helps keep them dependable, secure, and safe. You will own the reliability and operations side of the platform, ensuring model availability, performance, secure serving, and consistently safe generations for production traffic  will also work with a team that iterates quickly on new AI breakthroughs, with an emphasis on making speed reliable.

What you’ll be responsible for
  • Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale.
  • Build monitoring, alerting, and observability to detect ML-specific failures, output quality degradation, pipeline breakage, and model regressions before customers experience issues.
  • Harden model deployment workflows using canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely.
  • Drive the security posture for the model fleet with secure model serving, abuse and misuse detection,
    rate limiting
    , and protection against adversarial usage patterns.
  • Operationalize safety systems for generative media, including content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without sacrificing performance.
  • Lead incident response for model API outages and degradations, run postmortems, and drive engineering changes to prevent recurrence.
  • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic.
  • Partner with model and infrastructure teams so reliability, security, and safety requirements are built into how new models are onboarded to the platform.
What you’ll bring
  • 5+ years of professional experience, including 2 years operating production ML or high-scale API systems, ideally with on-call ownership.
  • Experience supporting diffusion models in production.
  • Strong systems fundamentals
    : distributed systems, networking, observability, and incident management.
  • Working knowledge of modern generative models (diffusion, transformers) and their production failure modes
    .
  • Familiarity with security and safety practices for ML systems, including abuse prevention or trust and safety engineering experience (strong plus).
  • A bias toward automation, measurement
    , and blameless postmortems
    .
Tools you’ll use
  • Python, torch, diffusers, Kubernetes, fal Python SDK
Role details
  • Department: EngineeringML
  • Employment type: Full time
  • Location: Remote in APAC (India, Australia, or New Zealand)

You’ll have access to fal’s massive GPU cluster for inference and evaluation, and you’ll work with a team focused on quickly iterating and deploying AI improvements while maintaining reliability.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary