Observability Engineer - Network Telemetry Specialist
Listed on 2026-07-24
-
Software Development
Backend Developer, Software Architect, Cloud Engineer - Software
Primary
Skills:
Java (Expert), Apache Kafka (Expert), PostgreSQL (Expert), Grafana (Expert), Telemetry / Network Observability (Expert)
Contract Type: W2 Only
Duration: 12+ Months
Location: Redwood City, CA (Hybrid, remote an option with occasional travel)
Pay Range: $90 - $95 on W2
Job SummaryWe are seeking a Senior Staff Engineer – NPE Observability to lead the architecture and evolution of a large-scale telemetry and observability platform. This role will drive the technical strategy for real-time data ingestion, streaming, and visualization across a global network infrastructure. The ideal candidate will possess deep expertise in Java, Kafka, PostgreSQL, Grafana, and distributed systems, with a proven track record of designing highly scalable, low-latency observability platforms.
Key Responsibilities- Architect and optimize scalable telemetry ingestion and storage platforms using Java and PostgreSQL.
- Design and enhance high-throughput Apache Kafka streaming pipelines for real‑time telemetry processing.
- Define enterprise observability standards and build advanced Grafana dashboards for monitoring global infrastructure.
- Architect stateful stream‑processing solutions using technologies such as Apache Flink.
- Evaluate and prototype emerging observability technologies including Model‑Driven Telemetry (MDT), Click House, and Thanos.
- Define platform architecture, technical standards, and long‑term engineering roadmaps.
- Establish and monitor SLIs/SLOs to ensure high availability, reliability, and platform performance.
- Lead complex root cause analysis, performance tuning, and architectural improvements for mission‑critical systems.
- Collaborate with software engineering, network engineering, and infrastructure teams to translate business requirements into scalable technical solutions.
- Mentor engineering teams and drive technical excellence through architecture reviews and best practices.
- 10+ years of software engineering experience with expertise in distributed systems.
- 5+ years of experience building large-scale network engineering, telemetry, or observability platforms.
- Expert‑level proficiency in Java backend development.
- Strong expertise with Apache Kafka, including cluster architecture, messaging, and stream processing.
- Advanced experience with PostgreSQL schema design, optimization, and performance tuning.
- Expert‑level experience developing enterprise dashboards using Grafana.
- Strong understanding of distributed systems, real‑time streaming, and high‑throughput data processing.
- Experience with Prometheus, Thanos, Click House, or similar observability platforms.
- Experience defining SLIs, SLOs, monitoring strategies, and incident management.
- Strong stakeholder management, technical leadership, and architectural design skills.
- Experience with Apache Flink or other stream‑processing frameworks.
- Knowledge of Model‑Driven Telemetry (MDT) and modern telemetry architectures.
- Experience working with cloud‑native observability platforms and Kubernetes environments.
- Familiarity with ITIL processes and enterprise operational best practices.
- Experience mentoring senior engineers and leading technical strategy across large organizations.
- Bachelor's or Master's degree in Computer Science, Software Engineering, or a related technical discipline.
- Experience designing globally distributed, high‑availability observability platforms.
- Strong background in Agile/Scrum software development methodologies.
- Proven ability to drive long‑term technical vision, innovation, and engineering excellence in enterprise‑scale environments.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).