Sr Monitoring & Observability Engineer, Los Angeles (On-Site
Listed on 2026-08-05
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support
Sr Monitoring & Observability Engineer, Los Angeles (On-Site)
Data Analysis Incorporated (DAI) is the controlling entity of the O'Neil family of businesses. DAI and its subsidiaries operate in diverse industries worldwide, including global equity markets, health care, financial services, digital news, and insurance. Our global footprint allows our teams to be responsive to customer needs in a timely and efficient manner. We are dedicated to using technology and innovation to bring change and growth to our businesses.
We believe in a dynamic workplace, creating engaging, informative products and services that help our customers succeed. Integrity is an essential characteristic for our firms and our associates; if this describes you, please apply!
The Senior Monitoring & Observability Engineer (TOC) is a senior-level infrastructure and reliability engineering role responsible for designing, implementing, optimizing, and supporting enterprise monitoring and observability platforms across networks, systems, cloud environments, and critical business applications. The position combines observability engineering, cloud and infrastructure operations, automation, and incident management responsibilities, including ownership of monitoring tools such as Datadog, Splunk, Solar Winds, Dynatrace, App Dynamics, Nagios, PRTG, and Zabbix.
Acting as a technical escalation point, the role partners closely with Infrastructure, Security, Dev Ops, and IT Operations teams to improve system reliability, alert quality, operational efficiency, and service availability while supporting SRE-aligned practices such as automation, root cause analysis, SLIs/SLOs, and continuous operational improvement.
Targeted pay range $115K - $125K Base pay, + 10% yearly bonus target.
Location:
11065 Beatrice St., Los Angeles, CA 900066
Duties and Responsibilities:
- Monitor and manage IT infrastructure, network systems, and business applications using enterprise monitoring tools, aligned with the TOC Sr. Engineer scope.
- Serve as the first point of escalation for TOC Engineers, providing advanced troubleshooting, guidance, and root cause analysis.
- Lead or support incident response, root cause analysis, escalation, and post-incident review processes; ensure issues are properly classified, escalated, and resolved efficiently.
- Take key roles in ITIL Incident, Problem, and Change Management processes.
- Build and tune monitoring and observability tooling — instrumentation, integrations, dashboards, alert logic, synthetic checks, log pipelines, and APM configuration — not just consume them.
- Develop and implement automation scripts and tooling to improve operational efficiency, alerting quality, and response times (Python, Power Shell, Bash, Ansible, or similar).
- Analyze system logs, network traffic, event data, and performance metrics to identify trends, reduce alert noise, and prevent outages.
- Document monitoring standards, troubleshooting steps, system configurations, dashboards, and runbooks for knowledge sharing.
- Collaborate with IT, Security, and Dev Ops teams to maintain system reliability and security posture.
- Work with vendors and service providers to resolve tool, platform, and infrastructure issues.
- Participate in 24/7 on-call rotations and provide leadership during major incidents, helping coordinate cross-functional resolution efforts.
- Mentor junior TOC/NOC engineers on monitoring tools, dashboards, alert handling, and incident response practices.
Qualifications & Requirements:
- Required Education, Experience, Certification/Licensure:
Bachelor's degree in IT, Computer Science, Networking, or a related field (or equivalent work experience). - 3+ years of experience in IT operations, network monitoring, or system administration, with hands-on experience implementing and tuning enterprise monitoring/observability platforms.
- Demonstrated experience building or implementing (not just using) one or more of:
Datadog, Dynatrace, App Dynamics, Splunk, Solar Winds Orion, Orion DPA, Nagios, PRTG, or Zabbix. - Advanced understanding of network protocols (TCP/IP, BGP, OSPF, VLANs, VPN, DNS, DHCP).
- Proficiency in Windows/Linux environments and at least one major cloud…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).