Lead Site Reliability Engineer; SRE
Juneau, Juneau Borough, Alaska, 99812, USA
Listed on 2026-07-18
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Summary
Lead Site Reliability Engineer (SRE) in the CDO Technology Group ensuring availability, latency, performance, and stability of critical infrastructure supporting data platforms, applications, and services.
Responsibilities- Availability
- Proactively monitor and identify potential issues impacting system availability.
- Implement automated alerts for outages or performance degradation.
- Collaborate with development teams to design and implement solutions that enhance resilience and reduce downtime.
- Latency
- Analyze performance metrics to identify and resolve latency bottlenecks.
- Implement performance optimization techniques and tools.
- Ensure new features and code changes do not introduce performance regressions.
- Performance
- Develop and maintain metrics dashboards for KPIs.
- Identify trends and anomalies indicating potential issues.
- Recommend and implement optimization strategies.
- Efficiency
- Optimize resource utilization and minimize unnecessary expenditure.
- Collaborate with development teams to optimize resource allocation for new applications.
- Release Management
- Participate in release planning to ensure smooth deployments.
- Develop automated deployment and rollback procedures.
- Monitor new releases and address issues promptly.
- Monitoring
- Design, implement, and maintain comprehensive monitoring infrastructure.
- Analyze data to identify potential issues and troubleshoot proactively.
- Develop alerts and notifications for critical events.
- Emergency Response
- Respond promptly to incidents and collaborate on resolution.
- Analyze root causes and implement preventive measures.
- Document incident responses and lessons learned.
- Participate in capacity planning to anticipate future workloads.
- Stay abreast of emerging technologies, trends, and best practices.
- Review architecture design for high availability and disaster recovery.
- Collaborate with reliability and infrastructure teams on observability, tracing, and alerting tooling.
- Bachelor's degree in Computer Science, Information Technology, or related field.
- 8+ years as a Site Reliability Engineer or equivalent.
- Proven experience monitoring, analyzing, and optimizing large‑scale distributed systems.
- Expertise in Linux systems administration, managing servers, OS, network configurations.
- Strong scripting and automation skills (Bash, Python, etc.).
- Familiarity with AWS.
- Experience with Dev Ops tools and practices (Git Lab CI/CD, Docker).
- Excellent troubleshooting and problem‑solving abilities.
- Strong communication skills for technical and non‑technical stakeholders.
- Passion for maintaining high availability, performance, and reliability in a fast‑paced financial environment.
- Competitive salary and comprehensive benefits package.
- Opportunity to work with cutting‑edge technologies and innovate solutions.
- Collaborative and supportive work environment with continuous learning.
- Competitive pay and bonuses, generous retirement plan, employee stock purchase plan with matching contributions.
- Flexible and remote work opportunities.
- Health care benefits (medical, dental, vision).
- Tuition assistance.
- Wellness programs (fitness reimbursement, Employee Assistance Program).
FINRA licenses are not required and will not be supported for this role.
Work FlexibilityEligible for remote work up to three days a week.
Commitment to Diversity, Equity, and InclusionWe strive for equity, equality, and opportunity for all associates, fostering an environment where people can bring their authentic selves to work and create belonging.
Equal Opportunity EmployerT. Rowe Price is an equal‑opportunity employer and values diversity of thought, gender, and race. We prohibit discrimination on the basis of race, religion, creed, color, national origin, sex, gender, age, disability, marital status, sexual orientation, gender identity or expression, citizenship status, military or veteran status, pregnancy, or any other classification protected by law.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).