×
Register Here to Apply for Jobs or Post Jobs. X

Platform ​/ Site Reliability Engineer

Job in New York City, Richmond County, New York, USA
Listing for: Sunset
Full Time position
Listed on 2026-08-20
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Job Description & How to Apply Below

Sunset Operations Engineer

Sunset operates customer-facing SaaS products, connector and ingestion services, asynchronous workers, high-volume data pipelines, model-backed systems, review tools, and customer-delivery paths. These workloads have different shapes, but they need a coherent foundation for infrastructure, delivery, observability, recovery, access, and cost.

You will build and operate the shared platform that lets our product, data, and AI teams ship reliable, secure, observable, and cost-aware systems without manual infrastructure work or operational risk growing linearly. You will write software and infrastructure, improve real engineering workflows, lead through incidents, and create paved roads teams can use without waiting on you.

This is not a deployment-operator or internal-IT role. Product, data, and ML teams remain responsible for the systems they build. You will give them the runtime, delivery, visibility, recovery, and operating patterns to own those systems well. You will partner closely with our Security Lead, but you will not be expected to run the entire security or compliance program.

Problems You Might Own

Create a small set of supported patterns for customer-facing services, connectors, scheduled jobs, data-processing pipelines, model-backed workloads, and evaluation runs. Define the contracts for environments, compute, state, networking, delivery, secrets, telemetry, failure handling, and recovery without forcing every workload into an inappropriate stack.

Make it straightforward for an engineer to create an environment, ship a safe change, understand a failed deploy or job, get the right access, recover a system, and know who owns the result. Build useful self-service and escape hatches while making unsupported paths and exceptions explicit.

Connect service, queue, job, pipeline, and model telemetry to the outcome that matters. Establish practical objectives, alerts, incident mechanics, replay and recovery paths, and reviews that remove recurring failure classes instead of only documenting them.

Expose cost and capacity in workload-relevant units, then improve them without hiding reliability, security, quality, or developer time. Work with Security to implement least privilege, secrets, logging, backup, deployment, and audit controls whose evidence comes from the systems that actually enforce them.

What You'll Do
  • Establish Sunset's current platform, workload, reliability, ownership, toil, recovery, cost, and technical-control baseline
  • Build reusable infrastructure-as-code modules, runtime templates, deployment workflows, environment contracts, and operational tooling
  • Create supported paths for customer-facing services, asynchronous and batch jobs, data pipelines, and model-backed workloads
  • Improve deploy safety, workload visibility, backup and recovery, incident response, replay, rollback, and durable remediation
  • Work with engineering teams to define useful service and pipeline objectives, ownership, escalation, and recovery paths
  • Build self-service for common infrastructure, environment, access, deploy, debugging, and recovery work without becoming a central approval queue
  • Make cloud and vendor cost understandable by service and workload and improve efficiency within explicit reliability and security bounds
  • Partner with Security on cloud identity, secrets, isolation, audit logging, vulnerability response, incident readiness, and automated control evidence
  • Support employees and contractors through bounded access, safe environments, release controls, documentation, and timely removal of authority
  • Use AI tools deeply for platform engineering and operations while verifying generated code, plans, queries, state changes, and incident conclusions
What Success Looks Like
  • Sunset's environments, runtimes, deploy paths, service and pipeline owners, reliability risks, recovery gaps, manual work, and infrastructure costs are visible and prioritized
  • One consequential failure or toil class is materially reduced in your first 90 days, and another team can use the resulting paved road without case-by-case help
  • Product, data, and AI teams can ship and understand their systems faster while retaining clear operating ownership
  • Priority services and pipelines have useful objectives, actionable telemetry, tested recovery paths, and incident learning that removes recurring failures
  • Common platform work becomes self-service while exceptions remain explicit, owned, monitored, and time-bounded
  • Cloud cost and capacity are understandable in workload-relevant units and improve without hidden reliability, security, or developer-time regressions
  • Security and customer-trust evidence becomes easier to produce because it reflects current technical controls
You Might Thrive Here If
  • You have personally owned production cloud infrastructure and delivery or reliability systems across multiple services, including an asynchronous, batch-data, or model-backed workload
  • You are a strong software engineer who is comfortable changing application,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary