Site Reliability Engineer
Listed on 2026-08-02
-
Software Development
AWS
Job Description
As a Senior Dev Ops / SRE Engineer on contract, you will be embedded with the Central Technology AI enablement team, working alongside engineers from Direct, Pitch Book, Retirement, and other business units. Your initial focus will be on the SRE and hosting side of our growing AI platform footprint. As the platform matures, we expect this role could extend into hands‑on contributions to the AI enablement components themselves, including the agent registry, agent runtime platform, and MCP tooling.
This is a hands‑on individual contributor role. You will have the opportunity to shape how a large financial services organization operates production AI infrastructure at scale, working with a small, senior team that moves quickly and makes decisions in the open.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances.
If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy:
- 5+ years of hands‑on Dev Ops, SRE, or platform engineering experience in a production environment.
- Strong Kubernetes experience — you have run production Kubernetes workloads and debugged real cluster issues.
- Solid experience with cloud infrastructure (AWS strongly preferred; Azure also relevant).
- Proficiency in Python, Bash, or a similar language for automation and tooling.
- Experience with modern CI/CD tooling (Git Hub Actions, Jenkins, Git Lab CI, or similar) and with Git Ops patterns.
- Strong debugging, systems‑thinking, and root‑cause analysis skills.
- Clear written and verbal communication — comfortable working across geographically distributed teams and with engineers outside your immediate area.
- Direct experience hosting or operating LLM gateways, model gateways, or LLM‑adjacent infrastructure (LiteLLM, model routers, inference platforms).
- Familiarity with the emerging MCP and agent ecosystem (MCP servers, agent runtimes such as KA Agent or Lang Smith Fleet, agent registries).
- Experience with observability platforms (i.e. Langsmith) and cost‑attribution patterns for shared multi‑tenant infrastructure.
- Experience operating platforms in a regulated or financial services environment.
- AWS certifications.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).