×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

Software Developer 4

Job in Menlo Park, San Mateo County, California, 94029, USA
Listing for: Stanford University
Full Time position
Listed on 2026-07-31
Job specializations:
  • Software Development
    Unix/Linux
Salary/Wage Range or Industry Benchmark: 150000 - 190000 USD Yearly USD 150000.00 190000.00 YEAR
Job Description & How to Apply Below

Job Description

The Rubin U.S. Data Facility at SLAC seeks a senior Software Developer / Site Reliability Engineer to lead and execute complex software and infrastructure work supporting Rubin Observatory production services. This role will help operate, automate, and improve the systems that support Alert Production, Data Release Production, database-backed services, Kubernetes applications, Weka and object-storage services, data movement, and other USDF-hosted operational platforms.

Job Description

The Rubin U.S. Data Facility at SLAC seeks a senior Software Developer / Site Reliability Engineer to lead and execute complex software and infrastructure work supporting Rubin Observatory production services. This role will help operate, automate, and improve the systems that support Alert Production, Data Release Production, database-backed services, Kubernetes applications, Weka and object-storage services, data movement, and other USDF-hosted operational platforms.

This position combines senior software development, site reliability engineering, production operations, deployment engineering, and cross-team technical leadership. The successful candidate will work across Rubin Data Management, SQuaRE, Prompt Processing, DRP, database, storage, networking, and USDF infrastructure teams to ensure that critical services can be deployed, monitored, debugged, upgraded, recovered, and scaled safely.

The role is aligned with the Software Developer 4 level: leading and executing difficult or complex programming and analysis work, contributing broad technical responsibility across multiple functions, and interfacing with other complex systems and programs.

Key Responsibilities
  • Lead and execute complex software, automation, and operational engineering projects for Rubin production services at USDF, including Alert Production, DRP, data access, databases, and service infrastructure.
  • Support and improve Kubernetes-hosted services deployed through Rubin’s Phalanx/Git Ops environment, including Helm configuration, secrets management, environment promotion, release coordination, and operational rollback.
  • Maintain and troubleshoot PostgreSQL-backed services, including backup and restore, performance, schema-change coordination, monitoring, and incident response.
  • Help operate and debug services that depend on Weka storage, S3 gateways, object storage access, shared file systems, Butler repositories, batch processing, and high-throughput data movement.
  • Develop strategies, methods, tools, and procedures that improve reliability, automation, reproducibility, observability, and operational efficiency across USDF production services, consistent with Software Developer 4 expectations for process improvement and innovation.
  • Direct or drive all phases of selected technical projects, including requirements analysis, design, implementation, testing, deployment, documentation, and operational handoff.
  • Lead testing, debugging, change control, reporting, and documentation for major service and infrastructure changes.
  • Build and maintain monitoring, logging, alerting, dashboards, runbooks, and operational documentation using tools such as Prometheus, Grafana, Loki, and related systems.
  • Provide subject-matter expertise for complex technical problems involving Kubernetes, databases, storage, networking, deployment automation, and distributed production services.
  • Recognize opportunities for high-impact, long-term reliability improvements that span team or organizational boundaries, and recommend or implement actions to resolve them.
  • Mentor technical staff and collaborate with developers, operators, database engineers, storage engineers, and scientists to improve service reliability and operability.
Desired Qualifications And Experience
  • Bachelor’s degree and ten years of relevant experience, or a combination of education and relevant experience, consistent with the Software Developer 4 profile.
  • Demonstrated experience designing, developing, testing, deploying, operating, and maintaining complex applications or infrastructure services.
  • Strong experience with Kubernetes, Helm, Git Ops, CI/CD, Linux systems, Python, shell scripting, and production service automation.
  • Strong understanding of relational databases, especially PostgreSQL, including operational monitoring, performance troubleshooting, backup/restore, and schema-change coordination.
  • Experience with large-scale storage or data systems, ideally including Weka, S3/object-storage interfaces, shared file systems, high-throughput data transfer, or scientific data repositories.
  • Experience with observability systems such as Prometheus, Grafana, Loki, alerting systems, logs, metrics, and operational dashboards.
  • Strong understanding of software development life cycle, quality-control practices, change management, incident response, and operational documentation.
  • Exceptional written and oral communication skills for technical and non-technical audiences, including the ability to work across distributed teams and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary