Staff AI Platform & Reliability Engineer
Listed on 2026-08-12
-
Software Development
Backend Developer, AI Engineer (Applied/Software), AWS, DevOps
Location
:
Santa Ana, California — on‑site
Employment
type
:
Full‑time
Salary
: $172,000 – $220,000
Created.ai is an AI‑powered creative platform combining text, image and video generation into production tools for commercial creative work. The platform is operated by eJam, a multi‑brand consumer products and technology company.
We are expanding our engineering team with two appointments covering our AI platform infrastructure and our product engineering surface. Both roles carry substantial technical ownership and report directly into engineering leadership.
This position is AI‑native by design. Our engineers work with agentic coding tools as a primary part of their daily workflow, which allows a small team to operate at a scope that would conventionally require a much larger one. We are looking for engineers whose judgement, architectural thinking and review discipline scale that leverage rather than being replaced by it.
Role summaryWe are seeking a Staff Engineer to take ownership of the AI generation services, provider integrations, billing integrity, tenant security and overall platform reliability underpinning created.ai. The role is Python and GCP‑first, with sufficient Type Script proficiency to trace and modify cross‑service contracts.
This is a senior individual contributor position with architectural authority over the platform layer.
Key responsibilities
- Own the design, delivery and operation of AI generation services and all third‑party provider integrations
- Ensure billing and usage‑metering integrity across the platform, including reconciliation and credit accounting
- Maintain tenant isolation and platform security controls
- Lead reliability engineering: failure handling, degradation strategy, capacity and incident response
- Strengthen deployment safeguards, observability and cost attribution
- Set technical standards for the Python services and mentor engineers working within them
- Python 3.12 with FastAPI, Pydantic, asyncio and strict typing in production
- Google Cloud Platform:
Cloud Run, Pub/Sub, Cloud Tasks, GCS, Firestore, Cloud SQL/Postgres - Demonstrated experience with webhooks, queues, retries, idempotency, dead‑letter queues and durable asynchronous jobs
- SQL Alchemy and Alembic, together with billing or usage‑metering experience
- Infrastructure and delivery tooling:
Terraform, IAM/OIDC, Secret Manager, CI/CD - Production integrations with image, video or LLM providers
- Working proficiency in NestJS/Type Script sufficient to modify cross‑service contracts
- Strong background in observability, cost tracking and incident debugging
- Fluency with agentic coding tools (Claude Code, Codex, Cursor or comparable), including the ability to scope work for them, review their output critically and maintain architectural coherence across AI‑assisted changes
- Health, Dental, Vision
- 401k Plan
- PTO Plan
- 14 observed local holidays
- Stock options
- Amazing, pet‑friendly office environment
- Equipment budget and learning allowance
- Real technical ownership within a small engineering team, and visible impact on a commercially active AI product with real users and real scale constraints
- Direct involvement in product decisions — the team is small enough that there is nowhere to hide, in both directions
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).