PlatformOps Engineer; Emerging AI
Listed on 2026-09-18
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
About DAT
DAT Freight & Analytics is an award-winning employer of choice and a next-generation SaaS technology company that has been at the leading edge of freight and logistics innovation for nearly five decades. Founded in 1978, DAT operates the largest freight marketplace in North America — processing 250 million+ load posts annually and maintaining one of the largest repositories of freight market transaction data in the world.
On a defined path to $1 billion in revenue, DAT deploys a suite of software solutions, machine learning models, and intelligent automation tools that help brokers, carriers, and shippers price freight accurately, source capacity, reduce risk, and operate more efficiently. With nearly 700 teammates across offices in Denver, CO;
Portland, OR;
Seattle, WA;
Springfield, MO;
Toronto, ON; and Bangalore, India, DAT combines the credibility of a multi-decade market leader with the drive of a company that is not done disrupting the industry it helped build.
For more information, visit
Job Final date to receive applications: 10/31/2026
The OpportunityAs a Staff AI Platform Engineer, you'll build and own the platform that every AI and machine learning workload at DAT runs on. Freight is an uncertain business, and the models we ship reduce that uncertainty: rate forecasts, load-to-truck matching, document extraction, fraud signals, and the agentic workflows our brokers and carriers use to move freight faster. None of that reaches a customer without a platform that makes training, serving, evaluating, and monitoring models routine instead of heroic.
You’ll set the architectural direction for that platform, from the model gateway and inference layer through feature and vector storage, evaluation harnesses, and production observability. This is a highly visible role for an engineer who wants their work multiplied across every AI team in the company.
What You'll Do- Platform Ownership: Design, build, and operate the shared services that engineers and business users use to ship models: a model gateway for LLM access, inference endpoints for real-time and batch scoring, feature storage, vector search, and a common SDK.
- Technical Leadership: Lead architecture for large-scale AI systems, write the design documents, and drive alignment across Product, and Engineering and Business users on how models get built and shipped at DAT.
- Cloud Architecture: Architect and run scalable, reliable AI infrastructure on AWS, including Bedrock, Sage Maker, EKS, Redpanda, MSK (Kafka), Lambda, and S3, all defined in Pulumi.
- Evaluation and Quality: Build the offline and online evaluation systems that tell us whether a model or prompt change is an improvement, including regression suites, LLM-as-judge pipelines, A/B and shadow testing, and drift detection.
- Cost and Performance: Own inference cost and latency as first-class metrics. Right-size GPU and serverless capacity, tune batching, caching, and quantization, and give teams clear visibility into what their workloads cost.
- Safety and Governance: Implement guardrails, prompt and output logging, PII handling, access controls, and model and dataset lineage so AI systems meet our security and customer data commitments.
- Best Practices: Drive the adoption of modern software development practices across AI work, including automated testing, code reviews, CI/CD pipelines, and infrastructure-as-code.
- Mentorship: Mentor engineers and data scientists on production ML and distributed systems, and raise the operational bar of every team that builds on the platform.
- Incident Management: Lead the response and resolution for complex production incidents involving AI services, perform root cause analysis, and implement preventative measures.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).