Principal, Software Engineer - Observability
Listed on 2026-07-11
-
Software Development
Cloud Engineer - Software, Software Architect, Software Engineer, DevOps
Position Summary
As an observability principal engineer, you will be a key researcher and technical lead expert in the architecture and development of cloud native observability designs, managed services, and real-time telemetry software systems. You will use your depth of engineering and experience to create visionary software architectures and telemetry systems to achieve an observability software product portfolio. Additionally, you will design, develop, and implement large‑scale distributed systems that process large volumes of data focusing on scalability, latency, and fault‑tolerance in every system built.
You must be able to effectively communicate and build collaboration at all areas and levels of the business and engineering. An ideal candidate will be adept at architecting large‑scale distributed systems and proficient in coding Java. You will also utilize multiple telemetry technologies such as data models, metric libraries, data logging, distributed tracing, datalakes, data correlation, rule‑based alerting engines, real‑time data streaming pipelines, TSDBs, and application performance management (APM).
Working in a cloud infrastructure ecosystem consisting of VMs, Kubernetes, and containers, you will create metric software designs and solutions enabling real‑time monitoring and alerting of system and application metrics. You will lead research initiatives for cloud native designs and implementation within public and private clouds and will utilize TSDBs and correlation and data fusion of multiple data types and heterogeneous data streams coupled with artificial intelligence (AI) and learned behaviors for anomaly detection and forward projections of system and application expected behaviors.
This role will involve collaboration with enterprise architects, product managers, data scientists, engineers, and business managers to bring telemetry R&D projects into production. You will use a combination of open source and COTS technologies to solve real‑time telemetry problems at an enterprise‑wide scale. In parallel, you will lead the design of new systems and the redesign of existing systems to meet business requirements, changing needs, and integration of state‑of‑the‑art technology.
You will be an evangelist for the observability foundation socialization technology designs and implementations to engineering and business customers.
Design and architect cloud‑native observability solutions, mandating both system‑wide telemetry pipelines and real‑time monitoring capabilities. Lead the technical vision and roadmap, socializing designs with internal and external stakeholders. Execute large‑scale distributed system projects, ensuring scalability, latency, and fault‑tolerance. Drive research into AI‑enhanced anomaly detection and predictive analytics. Mentor and guide engineering teams, fostering best practices and technical excellence.
Minimum Qualifications- BS/MS in Computer Science, Engineering, or equivalent, with 10+ years in software engineering, design, and architecture.
- Deep understanding of the Java language and associated frameworks, and previous development of Java applications, libraries, SDKs, or services.
- Strong architecture leadership with demonstrated enterprise‑level software implementations.
- Previous architectural leadership in research, evaluation, creation of software designs, and distributed software implementations in production.
- Experience with technical leadership, software roadmaps, research and development, new software initiatives, and customer and engineering coordination and engagement.
- Full‑stack cloud software development experience.
- Option 1:
Bachelor’s degree in computer science, computer engineering, software engineering, or related area and 5 years’ experience in software engineering or related area. - Option 2: 7 years’ experience in software engineering or related area.
- API/lib/SDK development, integration, and utilization.
- Cloud technologies and cloud native designs.
- Cloud infrastructures and technologies such as Open Stack, Azure, GCP, or AWS.
- Large‑scale distributed systems experience including scalability and fault…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).