Sr Manager IC, FinOps & AI Enablement — Analytics Team (Remote - Eligible
Newport News, Virginia, 23601, USA
Listed on 2026-08-16
-
Software Development
Data Engineering
Sr Manager IC, Fin Ops & AI Enablement
- Analytics Team (Remote
- Eligible)
The role Own the cost and AI-leverage layer of the analytics platform behind Capital One Shopping - the systems between petabyte-scale data and the humans and tools that query it. This is an own-and-build role, not a maintenance seat: you own the production platforms below, and in your first six months you ship three net-new systems on top of them. You'll report to the Engineering Director for the Shopping data platform as one of two senior IC pillars of the analytics org.
You own the technology choices and the strategic backlog and priorities for this surface - and you carry ongoing production support and an on-call rotation for these platforms and the models on them. You build the three net-new systems below on top of that operational base.
What you own- A data warehousing platform serving ~250K queries/day over ~20PB.
- An event ingestion pipeline taking in 6–7 billion events/day.
- The Airflow orchestration platform.
- A cost-attribution pipeline that parses Trino query logs at production traffic and attributes real AWS dollars to every report, query, user, and dbt model - reconciled against a ~$750K/month cloud bill.
- A forecast-driven autoscaling control loop for the shared Trino cluster and dbt worker pool - turning today's event-only Nomad autoscaler (Prime Day, Cyber Week) into steady-state, forecast-driven capacity - the single largest lever on the analytics AWS bill.
- A production Gen AI system - natural-language-to-SQL or RAG over the data catalog - with real LLM tool-use, grounding, and cost guardrails, adopted by internal teams.
The stack Kafka streaming backbone into an S3 lakehouse (Hive + Iceberg), Cassandra, Postgres, DynamoDB, Elastic Search, Aurora MySQL. Queried through Trino/Presto and Spark SQL, modeled in dbt, orchestrated on Airflow and Nomad and containers (Docker/Kubernetes), on a deep AWS footprint. SQL and Python daily;
Go, Java, and Type Script/JavaScript across the surrounding platform.
The day-to-day 'Manager' is the level, not the job - this is an individual-contributor role, and you'll spend most of your day hands-on in the editor. Roughly 70% building: writing the log-parsing and cost-attribution logic and its dbt models, building and tuning the forecast-driven autoscaler control loop, and building the RAG / natural-language-to-SQL system yourself. The other ~30% is technical coordination - reconciling your cost numbers with Finance, the R&D memo, aligning report owners - not status decks or people-management.
Daily rhythm is multi-terminal Claude Code: query-log analysis in one, dbt work in another, AI iteration in a third. No direct reports - you build.
- 10+ years engineering experience owning and supporting mission-critical applications and platforms in production
- Deep experience with Kafka and streaming technologies and platforms
- Experience with enterprise data technologies and platforms
- A track record working with extremely large traffic and data volumes
- Fluency across a real stack:
JavaScript, Java, HTML/CSS, Type Script, SQL, Python, and Go, open-source RDBMS and No
SQL databases, container orchestration (Docker and Kubernetes), and a broad range of AWS tools and services - Forecast- or workload-driven infrastructure scaling on a shared platform - you've scaled shared infra up and down against a forecast, not just event-driven bursts
- Production Gen AI - RAG design, vector/graph stores, LLM tool-use - on top of a longer ML/AI arc (ranking, recommendation, or comparable production ML predating the 2023 Gen AI boom)
- 0-to-1 delivery of a product that booked measurable revenue or adoption in its first weeks, plus a lead-engineer role modernizing an ingest pipeline off a legacy provider onto a cloud-native stack
- Experience directing a cross-team migration or deprecation at senior-executive scope to completion
- A broad background and a genuinely fast learner - the stack keeps moving
- Comfortable working with large teams on large-scale systems, and cross-functionally (weekly Finance/Biz Ops and Analytics cadences) Claude Code fluency - daily use, skill authoring, PR-level…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).