More jobs:
Job Description & How to Apply Below
Role Summary
NOC Lead will be responsible for leading enterprise-wide 24x7 monitoring and operational support across websites applications networks Windows servers cloud services and end-user computing platforms The role owns shift governance alert response major incident coordination vendor escalations ITIL process compliance BMC ticket governance SLA and KPI reporting and continuous service improvement The position requires hands-on command of engine-based monitoring products and Site
24x7 strong technical troubleshooting capability and disciplined leadership to maintain service availability performance reliability and operational excellence across the enterprise
- Command Centre Operations Lead 24x7 monitoring of enterprise websites applications APIs network devices Windows servers cloud services and other business-critical platforms
- Take operational ownership of engine-based products and monitoring tools including hands-on administration troubleshooting health checks and service restoration support
- Operate and optimize Site
24x7 monitoring for website and application availability response time transaction performance infrastructure health and alerting - Ensure monitoring thresholds probes synthetic checks dashboards notification rules and escalation paths remain accurate and aligned with business criticality
- Review monitoring coverage for new and changed services and ensure there are no unmanaged assets blind spots or unsupported alerts
- Reduce false positives and alert noise through regular tuning correlation suppression and improvement of monitoring logic
- Major Incident Management Ensure all alerts are validated prioritized acknowledged recorded and actioned within defined operational targets
- Lead the technical and operational response for Priority 1 and Priority 2 incidents including bridge coordination task allocation escalation and service restoration tracking
- Maintain clear communication with business stakeholders technical teams service owners and management throughout critical incidents
- Ensure incidents are linked to related problem change vendor and known-error records where applicable
- Drive post-incident reviews root cause analysis corrective actions and preventive measures for recurring or high-impact failures
- Verify that shift handovers include all active alerts open incidents pending vendor actions planned changes risks and follow-up items
- Shift Governance Supervise mentor coach and schedule NOC Engineers to ensure complete and effective coverage across all shifts including nights weekends and public holidays
- Prepare shift rosters manage leave coverage distribute workloads and maintain adequate staffing for operational and business requirements
- Set clear expectations for punctuality ownership ticket quality communication escalation discipline and a solution-oriented can-do attitude
- Conduct shift briefings knowledge-sharing sessions technical coaching performance reviews and competency development activities
- Maintain and enforce standard operating procedures runbooks escalation matrices checklists and shift handover standards
- Review team performance and take timely corrective action for process gaps missed alerts delayed escalations or recurring quality issues
- SLA Management Act as the primary operational liaison with third-party vendors managed service providers telecom providers application partners and support contractors
- Raise and track vendor cases provide required evidence and diagnostics coordinate troubleshooting sessions and elevate delays or service risks
- Monitor vendor response and resolution performance against contractual SLAs operational level agreements and service commitments
- Conduct regular vendor service reviews and follow up on chronic issues pending root cause reports recurring incidents and improvement actions
- Maintain vendor contact details support entitlements contract references escalation paths and service coverage information
- BMC Service Management Implement and enforce ITIL-based incident problem change event service request and knowledge management processes within NOC operations
- Use BMC Service Management BMC Helix BMC Remedy for ticket creation categorization assignment escalation tracking work notes resolution closure and reporting
- Ensure ticket records contain accurate timestamps impact and urgency troubleshooting evidence actions taken ownership customer communication and closure details
- Review aging breached reopened misclassified and unassigned tickets and drive timely corrective action
- Participate in change planning change advisory discussions maintenance windows implementation monitoring validation and rollback coordination
- Develop and maintain knowledge articles troubleshooting guides templates and known-error documentation
- Service Improvement Compile and share daily weekly and monthly reports covering incident trends alert volumes service availability SLA attainment…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×