Infrastructure Engineer
Listed on 2026-07-31
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Looking for a motivated and experienced Dev Ops / Site Reliability Engineer (SRE) to own and architect scalable infrastructure in a fast-paced, technologically innovative environment within a leading investment and data analytics firm. This is a key role reporting directly to executive leadership, perfect for a builder who thrives on end-to-end ownership and driving infrastructure strategy.
Role OverviewThis role involves designing, deploying, and maintaining robust cloud and on-premises infrastructure, ensuring high availability, security, and scalability of the core platform supporting AI, data processing, and financial analytics. The successful candidate will serve as the sole Dev Ops/SRE professional, making strategic decisions that influence platform reliability and operational excellence across the organization.
Key Responsibilities- Lead architecture design and implementation for cloud-native infrastructure, including virtual networks, container orchestration (Kubernetes, Docker), and cloud resources on AWS, Azure, or GCP.
- Develop, maintain, and optimize CI/CD pipelines, release workflows, and version control management to ensure smooth deployment processes.
- Manage container registries, secrets management (Key Vaults), storage solutions, and network configurations for secure and efficient system operation.
- Implement monitoring, observability, and alerting solutions (Prometheus, Grafana, ELK Stack, Data Dog, or similar) for comprehensive system health tracking and incident response.
- Collaborate with engineering teams to support scalable, reliable systems, integrating infrastructure with AI and data workflows as needed.
- Provide strategic input into infrastructure roadmap, automation, and security best practices, ensuring compliance and operational resilience.
- Own and troubleshoot production incidents, ensuring rapid resolutions and continuous system improvements.
- 3+ years of proven experience in Dev Ops, Site Reliability Engineering, or infrastructure engineering with a track record of building, deploying, and maintaining production-scale systems.
- Strong expertise in cloud infrastructure management (AWS, Azure, GCP), containerization (Docker, Kubernetes), and orchestration.
- Extensive hands-on experience with CI/CD pipelines, version control (Git), and automation tools (Jenkins, Git Lab CI, CircleCI, or similar).
- Proficiency in monitoring, observability, and incident management tools (Prometheus, Grafana, Data Dog, ELK, Splunk).
- Excellent communication skills with the ability to articulate technical options and system architecture to executive stakeholders.
- Ability to work independently as the sole Dev Ops/SRE resource, making strategic decisions and leading infrastructure initiatives.
- Experience in financial services, hedge funds, trading platforms, or fintech environments.
- Familiarity with infrastructure for AI and machine learning, including model deployment, MLOps pipelines, and AI workload orchestration (Lang Chain, MCP, MLflow, Langfuse).
- Hands-on experience with private package repositories (Artifactory, PyPI).
- Knowledge of security best practices, compliance standards, and non-compete considerations in finance.
- Prior experience at high-profile tech or quant firms such as Two Sigma, Renaissance Technologies, or similar.
- Infrastructure as Code:
Terraform, Cloud Formation, ARM templates - Orchestration & workflow automation:
Prefect, Airflow, Dagster
This position offers the opportunity to redefine infrastructure standards within a cutting-edge data-driven organization, impacting investment decisions and high-stakes financial systems.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).