×
Register Here to Apply for Jobs or Post Jobs. X

Senior Infrastructure & Reliability Engineer – AI Platform

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Writer
Full Time position
Listed on 2026-05-29
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 157700 - 277800 USD Yearly USD 157700.00 277800.00 YEAR
Job Description & How to Apply Below

Location

New York City, NY

Employment Type

Full time

Location Type

Hybrid

Department

Engineering, product & design

Compensation

  • SF & NYCBase Compensation $157.7K – $277.8K
    • Offers Equity

WRITER is committed to transparent, market-based compensation practices. Compensation offered will be determined by multiple factors such as role scope and complexity, location, experience, knowledge, and skills. Cash compensation is only one part of WRITER’s competitive Total Rewards package, which can also include equity, an immersive purpose-driven culture, career development, and thoughtfully designed benefits and well-being offerings.

🚀 About WRITER

WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs.

Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI.

Founded in 2020 with office hubs in San Francisco, New York City, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI.

📐 About the role

At WRITER, our mission to expand human capacity with super intelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI.

You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.

This is a hybrid position, based out of our New York City hub. You'll report to our director of engineering.

🦸🏻♀️ What you'll do

  • Use and build AI native approaches for operational tasks and infrastructure management and platforms using Python, Go, or similar languages, significantly reducing manual toil across our production environment

  • Design and implement scalable, fault-tolerant infrastructure AI solutions on public cloud providers (AWS, GCP, Azure) to support WRITER's rapidly expanding, high-traffic AI platform

  • Own the reliability, performance, and efficiency of WRITER’s core services, defining and upholding stringent Service Level Objectives (SLOs) and Error Budgets

  • Own the observability stack for monitoring, logging, and alerting systems to ensure rapid detection of issues across our complex distributed systems

  • Lead incident response, post-mortems, and root cause analyses, applying learnings to proactively prevent future outages and build a more resilient system architecture

  • Collaborate closely with product and engineering teams, providing expert guidance on system design for reliability, performance, and scalability from conception through launch

⭐️ What you need

  • A solid 7+ years of experience in Infrastructure engineering, Dev Ops, Production engineering, Cloud platform or a similar role focused on building and operating large-scale, high-availability production systems

  • Deep expertise with cloud platforms (AWS strongly preferred), containerization technologies like Docker and Kubernetes, and Infrastructure-as-Code tools such as Terraform

  • Strong proficiency in programming languages such as Python, Java, Go for automation and monitoring

  • Knowledge of monitoring and logging tools (e.g.,…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary