AI Infrastructure Engineer
Miami, Miami-Dade County, Florida, 33222, USA
Listed on 2026-07-14
-
Software Development
Backend Developer, DevOps, AI Engineer (Applied/Software), Cloud Engineer - Software
Magen Financial LLC is a fintech brokerage company focused on delivering modern, technology‑driven financial services and trading solutions. The firm combines brokerage expertise with advanced data, AI, and engineering capabilities to support informed decision‑making and efficient operations. We are building secure AI and data systems for financial markets workflows, operating primarily on the Microsoft Azure stack and integrating closely with .NET‑based
internal systems. Because we work with sensitive financial data, we care deeply about security, auditability, correctness, and reliability. The company values innovation, reliability, and data integrity as core components of its service offering.
This is a full‑time, hybrid AI Infrastructure Engineer role based in Miami, FL, with flexibility for partial work from home. You will own the platform layer for our AI services, making advanced models reliable, secure, and usable in production. Responsibilities include designing and operating infrastructure for serving large and smaller specialist models; building secure internal APIs for AI‑powered applications; deploying and managing GPU‑based workloads across development and production;
and integrating model serving with retrieval systems, databases, internal services, and authentication. You will build CI/CD pipelines for model and application deployments, implement observability for latency, throughput, errors, GPU utilization, and service health, and support batch, interactive, and evaluation inference workloads. Working closely with AI engineers, you will support fine‑tuning, evaluation, and deployment while implementing secure access controls, audit logging, and environment separation.
This is a hands‑on platform role focused on turning AI models into dependable internal services, optimizing reliability, cost, and performance, and defining standards for model packaging, versioning, rollout, and rollback.
- Designing and operating infrastructure for serving large and smaller specialist models
- Building secure internal APIs for AI-powered applications
- Deploying and managing GPU‑based workloads across development and production
- Integrating model serving with retrieval systems, databases, internal services, and authentication
- Building CI/CD pipelines for model and application deployments
- Implementing observability for latency, throughput, errors, GPU utilization, and service health
- Supporting batch, interactive, and evaluation inference workloads
- Supporting fine‑tuning, evaluation, and deployment while implementing secure access controls, audit logging, and environment separation
- Optimizing reliability, cost, and performance
- Defining standards for model packaging, versioning, rollout, and rollback
- 5+ years of infrastructure, platform engineering, Dev Ops, SRE, or ML platform experience
- Strong experience with Microsoft Azure
- Strong experience with Kubernetes, preferably AKS, and with Docker and containerized services
- Strong C# / .NET ecosystem experience, especially for internal service integration
- Good Python skills for automation, AI infrastructure, and scripting
- Experience building production APIs and internal developer platforms
- Experience with CI/CD, infrastructure as code, and observability
- Experience operating systems with high security and reliability requirements
- Ability to debug complex issues across applications, infrastructure, networking, and storage
- Experience with GPU clusters or distributed compute, and serving large language models or other deep learning models
- Experience with high‑throughput batch processing
- Experience with model registries, MLflow, or Azure Machine Learning
- Experience with financial services infrastructure or secure internal platforms in regulated environments
- Experience with retrieval‑augmented generation systems, vector databases, or search infrastructure
- Experience with performance tuning for latency‑sensitive services
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).