Software Engineer - AI/ML Platform
Listed on 2026-10-07
-
Software Development
AI Engineer (Applied/Software)
Fanatics Commerce is the global leader in licensed sports merchandise, operating a vertically integrated platform that designs, manufactures, and delivers officially licensed apparel, jerseys, headwear, and collectibles for major leagues, teams, and events worldwide. With more than 900 e-commerce sites and a global omnichannel presence across digital, in-venue, and retail, Fanatics Commerce reaches fans in over 180 countries and powers official fan experiences for many of the world’s most iconic sports properties.
At Fanatics, we bring our BOLD Leadership Principles to life every day - building championship teams, obsessing over fans, acting with entrepreneurial speed, and delivering with a determined and relentless mindset.
About The TeamThe AI/ML Platform pod builds Fanatics' internal AI platform: the paved road for every team building with LLMs, MCPs, and agents. Gateway, agent runtime, registry, retrieval, evaluation, and cost analytics, all running on a petabyte-scale lakehouse with enterprise-grade governance. You’ll join a pod that owns this platform end to end. This is a Staff-level role, and a hands-on one. You will choose the architecture, define the standards other teams build to, and stay in the code, shipping the hardest parts yourself.
The decisions you make are ones the organization lives with for years.
- Own the end-to-end architecture of the AI control plane: gateway, agent runtime,registry, retrieval, evaluation, and cost. Set technical direction, sequence the roadmap,and make the build-versus-buy calls.
- Build the enterprise LLM gateway: multi-provider routing, failover, caching, rate limiting,token optimization, and key management, with a path to self-hosted open-weight models where cost, latency, or data residency call for it.
- Build the Agent/MCP ecosystem on a managed runtime such as AWS Bedrock Agent Core : enterprise systems exposed as governed tools over MCP, a control-plane registry for discovery, ownership, versioning, and lifecycle, and a no-code path from prototype to governed production agent, surfaced through a company-wide chat portal.
- Own the unified retrieval layer and the data platform behind it: the governed front door plus the ingestion, chunking, embedding, and indexing pipelines and the vector and search backends that feed it, kept reproducible, observable, and quality-gated at petabyte scale.
- Build the evaluation and feedback loop: offline evals, tracing, and regression gates in CI/CD, with production failures and user edits turning into new eval cases and root-cause attribution across prompts, retrieval, tools, and models.
- Own governance and cost: RBAC, agent identity, least-privilege access to enterprise data, audit logging, use-case-level cost attribution, and the enterprise AI tool portfolio, standing up to finance, executive, and EU AI Act scrutiny.
- Raise the bar: set the engineering standards, mentor engineers, and represent the architecture and its trade-offs directly to senior leadership
- Build the evaluation and feedback loop: offline evals, tracing, and regression gates in CI/CD, with production failures and user edits turning into new eval cases and root-cause attribution across prompts, retrieval, tools, and models.
- Own governance and cost: RBAC, agent identity, least-privilege access to enterprise data, audit logging, use-case-level cost attribution, and the enterprise AI tool portfolio, standing up to finance, executive, and EU AI Act scrutiny.
- Raise the bar: set the engineering standards, mentor engineers, and represent the architecture and its trade-offs directly to senior leadership
- 8 to 12 years building production software, including 2+ years on AI/ML platform. Bachelor's or Master's in Computer Science, Engineering, or a related field.
- Staff-level impact: system design across multiple services, decisions that span teams, and platforms other engineers build on.
- Strong Python and FastAPI, plus one of Java, GoLang, Scala, or Type Script. Hands-on with Terraform, Kubernetes, Airflow, Postgres, Grafana, and Prometheus, and familiar enough with the AWS data science toolkit and libraries like PyTorch to partner credibly with data scientists. Chat UI or internal developer portal work is welcome.
enough with the AWS data science toolkit and libraries like PyTorch to partner credibly with data scientists. Chat UI or internal developer portal work is welcome.
- Hands‑on with LLMs and generative AI on Bedrock or Vertex: model routing, RAG, tool calling,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).