Software Engineer II - Recommendations
Listed on 2026-09-09
-
Software Development
Backend Developer, Cloud Engineer - Software, AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Software Engineer II - Recommendations(Boston, MA onsite 5x a week)
Why you should join the Recommendations Platform Team
The Recommendations Platform Team is responsible for developing and deploying machine learning-based recommendation systems at scale, and building out the foundation for new use cases for technologies such as embedding-based similarity search to power agentic workflows. We are evolving to act as a layer of product intelligence, making sense of customer data in order to personalize messages with the right item at the right time across multiple channels.
Our systems span large-scale data pipelines and querying workflows, batch training and inference, and low-latency online retrieval and ranking systems. We are also building the experimentation, tracking, and measurement capabilities needed to evaluate recommendation quality and business impact over time.
- Contribute to the architecture and evolution of backend services that power product recommendations across Klaviyo experiences (email, SMS, KAgent, onsite, etc.), meeting standards for reliability, performance, and clear APIs.
- Contribute to and maintain robust, large-scale data processing pipelines (e.g., using Apache Spark or similar frameworks) that transform raw events and catalog data into high-quality features and inputs for recommendation models, ensuring data quality and lineage.
- Collaborate closely with ML engineers and product stakeholders to product ionize recommendation models—defining high-level interfaces, feature contracts, and deployment patterns for batch and/or real-time inference systems.
- Contribute to the development of the vector database that powers recommendation, semantic search, and agentic use cases.
- Ensure data and service observability (metrics, logging, tracing, dashboards) to facilitate recommendations that are correct, explainable, fast, and highly available for all customers.
- Work with Product to break down projects into clear milestones, balancing the need for rapid experimentation with technical soundness and long-term maintainability.
- Lead data-driven decision making and A/B testing efforts—ensuring recommendation systems are instrumented with the right metrics, and independently interpreting results to guide future product and engineering iterations.
- Participate in on-call and incident response for the systems you own, driving major post-incident follow-ups that substantially improve the resilience and operability of our recommendation stack.
- Integrate AI into your and the team’s development workflow from the ground up—for example, using AI to accelerate development, automate complex tests, or build smarter monitoring and debugging tools.
- Share knowledge, mentor junior engineers, and define best practices on working with large-scale data frameworks, distributed systems, and integrating ML into production systems.
- 2+ years of professional software engineering experience with a focus on backend and distributed systems at scale; you have a proven track record working on production services and optimizing for latency, reliability, and operability as well as business requirements.
- Proficient in Python and open to working in other languages
- Comfortable with cloud-native architectures (AWS preferred) and container orchestration (e.g., Kubernetes); you manage infrastructure and CI/CD pipelines as a core part of your development process.
- Experience in data-driven decision making and A/B testing—you can define (or are interested in learning how to) how to instrument experiments, read and interpret results, and ensure learnings are folded back into system design.
- Comfortable designing and querying data models in relational, analytical, and No
SQL data stores (e.g., Postgres, MySQL, data warehouses, Redis, vector databases). - Feel at home with modern Dev Ops practices (CI/CD, monitoring, alerting) and how to apply them to architect large-scale data and recommendation systems.
- Track record of owning features end-to-end—from initial technical design and implementation through rollout, monitoring and sustained iteration.
- Excellent technical collaborator and communicator: you can…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).