Platform Reliability Engineer, Associate/VP - Aladdin
Listed on 2026-09-10
-
Software Development
AI Engineer (Applied/Software), AWS
hackajob is collaborating with Black Rock to connect them with exceptional professionals for this role.
About this role
Your Team
The Production Engineering team within Aladdin Engineering delivers the secure, resilient, and scalable platform capabilities that power critical engineering workflows across the Aladdin ecosystem. We are at the forefront of modern platform operations, combining automation, observability, and applied artificial intelligence to improve how engineers build, operate, and support technology AI becomes increasingly embedded in engineering practices, our team is shaping the next generation of operational excellence through intelligent tooling, self-service experiences, and secure platform services that enhance reliability, accelerate innovation, and improve developer productivity.
Your Role and Impact
As an AI Platform Reliability Engineer, you will help define the future of AI-enabled platform operations within Aladdin. You will design and operate highly reliable services, develop intelligent automation, and build AI-assisted capabilities that make complex systems easier to manage and support. Working at the intersection of platform engineering, secrets management, observability, and applied artificial intelligence, you will deliver solutions that reduce operational overhead, improve system resilience, accelerate incident response, and empower engineers with faster access to knowledge and insights.
Your work will directly influence platform reliability and the engineering experience across a broad technology landscape.
Your Responsibilities
In every role at Black Rock, you'll be expected to apply sound judgement and critical thinking to solve complex problems, adapt as the business evolves, and embrace the curiosity to explore new approaches and technologies while challenging outcomes. The scope of this role also includes the following responsibilities:
- Design, build, and operate reliable platform services through automation, infrastructure-as-code, continuous delivery practices, Python-based workflow automation, and secure secrets-management lifecycle capabilities.
- Develop and operate AI-assisted platform capabilities, including retrieval-augmented knowledge solutions, operational diagnostics, engineer self-service experiences, and the supporting data ingestion and retrieval pipelines.
- Monitor and analyze system health, performance, and capacity using observability tools and perform root-cause analysis for platform incidents.
- Participate in on-call rotations supporting production systems and drive continuous improvement through automation and AI-driven reduction of operational toil.
- Support disaster recovery planning and execution, application onboarding, platform upgrades, and change management for services using centralized secrets and secure configuration capabilities.
- Ensure automation and AI-enabled solutions are secure, explainable, production-ready, and aligned with enterprise and regulated-environment requirements.
Your Experience
- Bachelor's degree (or equivalent) in Computer Science, Engineering, Mathematics, or a related field.
- Strong Python development experience for automation, systems analysis, and integration with AI tooling.
- Experience with cloud platforms, including Amazon Web Services (AWS) and Microsoft Azure.
- 4+ years of Hands-on experience with infrastructure-as-code and configuration-as-code practices.
- Experience using monitoring, observability, and operational analytics tools.
- Foundational understanding of Artificial Intelligence (AI) and machine learning concepts, particularly within automation, observability, or developer productivity domains.
- Experience working with Linux-based environments, enterprise-grade container orchestration platforms supporting declarative infrastructure and horizontal scaling, retrieval-augmented generation solutions, vector databases, AI assistant frameworks, high-availability architectures, or large-scale technology environments.
Preferred Qualifications
- Experience with in financial services or other highly regulated technology environments.
- Experience building, operating, or supporting AI-enabled internal developer platforms and engineering productivity solutions.
- Certification or practical experience with enterprise secrets-management platforms.
- Relevant cloud certifications are beneficial.
For AT8 - 725 Ponce de Leon Ave NE, Atlanta Only the salary range for this position is USD$ - USD$ . Additionally, employees are eligible for an annual discretionary bonus, and benefits including healthcare, leave benefits, and retirement benefits. Black Rock operates a pay-for-performance compensation philosophy and your total compensation may vary based on role, location, and firm, department and individual performance.
Our benefits
To help you stay energized, engaged and inspired, we offer a wide range of benefits including a strong retirement plan, tuition reimbursement, comprehensive healthcare, support for working parents and Flexible Time Off (FTO) so you can…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).