×
Register Here to Apply for Jobs or Post Jobs. X

Lead Software Platform Engineer, MLOps

Job in Cambridge, Middlesex County, Massachusetts, 02140, USA
Listing for: TetraScience
Full Time position
Listed on 2026-08-13
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 200000 - 270000 USD Yearly USD 200000.00 270000.00 YEAR
Job Description & How to Apply Below

Who We Are

Tetra Science is the Scientific Data and AI Cloud company. We are catalyzing the Scientific AI revolution by designing and industrializing AI-native scientific data sets, which we bring to life in a growing suite of next gen lab data management solutions, scientific use cases, and AI-enabled outcomes.

Tetra Science is the category leader in this vital new market, generating more revenue than all other companies in the aggregate. In the last year alone, the world's dominant players in compute, cloud, data, and AI infrastructure have converged on Tetra Science as the de facto standard, entering into co-innovation and go-to-market partnerships.

In connection with your candidacy, you will be asked to carefully review the Tetra Way letter, authored directly by Patrick Grady, our co-founder and CEO. This letter is designed to assist you in better understanding whether Tetra Science is the right fit for you from a values and ethos perspective.

It is impossible to overstate the importance of this document and you are encouraged to take it literally and reflect on whether you are aligned with our unique approach to company and team building. If you join us, you will be expected to embody its contents each day.

The Role

We're looking for a Lead Software Platform Engineer working at the intersection of distributed systems and MLOps. You will own and scale our AI and data infrastructure as a product our customers build their science on, and that every other engineering team builds against. You will architect the cloud-based services and MLOps infrastructure that enable production-grade AI/ML workflows, working closely with Applied AI engineers, data engineers, and platform teams, and you will act as the technical design authority for how models, LLMs, and agents run in production.

The work spans the model and prompt lifecycle, the evaluation and observability harness, the security and tenant boundaries around prompts and retrieval, and the cost and latency controls that keep production AI economically viable at scale.

This is a highly impactful seat for an engineer who has shipped AI/ML infrastructure as a multi-tenant product rather than internal tooling. The model serving, MCP, and agent capabilities you build are consumed directly by scientists at the world's largest pharmaceutical companies, running against their own data under their compliance obligations. If that scope appeals to you, and you thrive on turning ambitious scalability and cost targets into concrete technical strategy inside a regulated environment, we'd love to talk to you.

What

You Will Do
  • Own the technical architecture of the AI/ML platform: the service and API surface our customers use to run models and agents against their own scientific data, and that Applied AI and data engineering teams build against internally.
  • Own the end-to-end model and prompt lifecycle across Databricks MLflow and AWS Bedrock, including registration, versioning, asset bundles, staged promotion, rollback, and multi-model serving.
  • Design the inference substrate for both real-time and batch AI workloads, including routing, batching, caching, concurrency control, GPU and accelerator capacity planning, handling of large binary inputs such as instrument images, and graceful degradation under load.
  • Integrate AI models and large language models (LLMs) into production systems using architectures like retrieval-augmented generation (RAG), and architect the agentic layer above them: tool and function calling, MCP-based tooling, and agent runtimes, deciding what belongs in the platform versus in the applications built on top of it.
  • Design security into the AI platform rather than leaving it to the applications above it, partnering with our security team on guardrails, prompt-injection and tool-abuse defenses, PII and PHI handling, and hard tenant data boundaries across prompts, retrieval, and tool calls.
  • Build the evaluation and quality infrastructure that makes AI shippable: offline and online eval harnesses, golden datasets, regression gates in CI, A/B and shadow deployment, and drift and hallucination detection in production.
  • Establish observability for the AI platform, including monitoring, alerting, logging, and distributed tracing, and set the SLI, SLO, and SLA model for systems whose outputs are probabilistic.
  • Design for reproducibility and lineage required in a validated pharma environment, with versioned data, code, prompts, and model artifacts, and an audit trail that can withstand customer and regulatory scrutiny.
  • Contribute to the infrastructure-as-code and deployment automation for the AI platform (Cloud Formation, AWS CDK), partnering with the team that owns production deployments to support multi-tenant infrastructure, online upgrades, and on-demand compute allocation.
  • Own production readiness for the AI platform with Applied AI engineers, data engineers, and platform teams: the performance, reliability, and cost-efficiency of models in production, plus incident response…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary