Job Description
Role:
Observability Engineer (Dynatrace, Splunk, Solar Winds, Service Now)
Duration: 12 Months
Role Description• Responsible for administration, support, optimization, and continuous improvement of enterprise monitoring and observability platforms including Solar Winds Orion, Splunk, and Dynatrace.
• Focus on monitoring infrastructure, applications, cloud environments, and enterprise services.
• Integrate monitoring solutions with Service Now for incident and event management.
Required Experience• 8+ years of experience in Infrastructure Operations, Monitoring Services, Enterprise Monitoring, or Observability platforms.
• 5+ years of hands-on experience administering enterprise monitoring tools.
• Strong troubleshooting, analytical, and stakeholder management skills.
• Experience supporting production infrastructure, cloud services, applications, networks, storage, and enterprise platforms.
• Hands-on experience with automation using Power Shell, Python, REST APIs, and operational workflows.
Essential SkillsSolar Winds Orion
• Network Performance Monitor (NPM)
• Server & Application Monitor (SAM)
• Alerting
• Custom Polling
• Dashboards
• Reporting
• Infrastructure Monitoring
• Network Monitoring
Splunk
• SPL Queries
• Data Onboarding
• Dashboards
• Reports
• Alerts
• Log Analytics
• Event Correlation
Dynatrace
• One Agent
• Smartscape
• Distributed Tracing
• Synthetic Monitoring
• Real User Monitoring (RUM)
• Service Flow Analysis
Service Now
• Service Now Incident Management
• Event Management
• ITOM
• Automated incident generation
• Operational response workflows
Key ResponsibilitiesMonitoring Platform Administration
• Administer, support, and optimize Solar Winds Orion, Splunk, and Dynatrace platforms.
• Ensure platform availability, scalability, performance, and operational health.
• Perform upgrades, maintenance, configuration reviews, and platform optimization activities.
Monitoring & Observability Engineering
• Design and maintain monitoring solutions for:
• Infrastructure
• Network
• Cloud
• Storage
• Backup
• Applications
• Enterprise Services
• Implement proactive monitoring and observability capabilities across enterprise environments.
• Enhance monitoring coverage and improve operational visibility.
Alerting, Dashboards & Event Correlation
• Build dashboards, alerts, reports, service health views, and event correlation rules.
• Drive alert quality improvements and reduce alert noise across monitoring platforms.
• Establish monitoring standards and operational dashboards for business and technical stakeholders.
Service Now Integration
• Integrate monitoring platforms with Service Now Incident Management and Event Management workflows.
• Improve operational visibility through integrated monitoring and ITSM processes.
Incident Management & RCA Support
• Support Major Incident investigations and monitoring-related root cause analysis activities.
• Identify monitoring gaps and implement remediation plans.
• Contribute to service reliability and operational resilience initiatives.
Automation & Continuous Improvement
• Develop automation solutions using scripting and APIs to improve operational efficiency.
• Drive AIOps, observability, and event management maturity initiatives.
• Continuously evaluate monitoring effectiveness and recommend improvements.
RequirementsSailpoint
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: