Cloud Monitoring Engineer; Remote
Chesapeake, Virginia, 23320, USA
Listed on 2026-07-08
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Exclusion States
AK, CA, CO, CT, DC, HI, LA, MA, MN, MO, NE, NV, NH, NJ, NM, NY, ND, OR, PR, RI, VT, WA, WY.
Future NeedActively Interviewing
LocationRemote in any United States jurisdiction not excluded from this job advertisement.
Job SummaryProvide the eyes‑on‑glass excellence a mission‑critical Department of Veterans Affairs (VA) platform demands. As a Cloud Monitoring Engineer, you will build, tune, and maintain the observability stack tracking latency, error rate, saturation, volume, and incident‑free availability across 300+ applications.
Position DescriptionThe Cloud Monitoring Engineer builds and maintains the Capabilities and Services Dashboard and supports monitoring infrastructure, ensuring automated alerting detects production issues before user‑reported tickets arrive.
Minimum/General Experience5 years of experience in cloud monitoring and observability engineering
Minimum EducationBachelor's Degree in computer science, information technology, or related field;
Dynatrace Associate certification or equivalent observability platform certification (preferred)
- Excellent experience building and maintaining dashboards displaying latency, error rate, saturation, volume, and incident‑free availability in real‑time (e.g., Dynatrace, Splunk).
- Excellent knowledge of the four Golden Signals.
- Excellent ability to individually monitor and track latency, error rate, saturation, volume, and incident‑free time.
- Excellent ability to implement dependency tracking within monitoring dashboards including latency, error rates, and transaction volumes.
- Excellent experience configuring automated alerts reflecting meaningful degradation or disruption while minimizing false positives.
- Excellent ability to maintain an accurate, complete, and auditable log of all alerts including alerted system, cause, timestamps, corrective actions, and responsible system.
- Above average experience supporting 24/7 monitoring operations and coordinating with on‑call Site Reliability Engineers (SREs) during active incidents.
- Above average knowledge of AWS Cloud Watch and integration with third‑party observability tools in a Gov Cloud environment.
- Experience supporting federal government programs and enterprise‑scale applications operating in cloud‑based or hybrid environments.
- Excellent verbal and written communication skills.
- Assignment Location
- Remote. - Sedentary Work
- Exerting up to 10 pounds of force occasionally and/or a negligible amount of force frequently or constantly to lift, carry, push, pull or otherwise move objects. - Typing, communicating, repetitive motions.
- Close visual acuity to prepare and analyze data, view computer monitors and read. May need to view presentation screens and other visual aids in a virtual setting.
- Inside environmental conditions with protection from outside elements.
Active Federal Civilian Public Trust clearance required.
- U.S. Citizenship or Permanent Resident that has lived in the United States for at least 3 years.
- Build and maintain the Capabilities and Services Dashboard displaying real‑time latency, error rate, saturation, volume, and incident‑free availability for all capabilities and services.
- Implement dependency tracking within the dashboard including latency, error rates, and transaction volumes for all capability and service dependencies.
- Configure and tune automated monitoring and alerting mechanisms ensuring personnel are alerted to production issues prior to receipt of user‑reported tickets.
- Ensure all capabilities and services are individually monitored and tracked for latency, error rate, saturation, volume, and incident‑free time.
- Maintain an accurate, complete, and auditable alert log including alerted system, description, timestamps, corrective actions, and responsible system.
- Continuously tune alert thresholds to reflect meaningful degradation or disruption while minimizing false positives.
- Support the Capabilities and Services Monitoring Plan defining alert conditions, thresholds, notification mechanisms, and escalation paths.
- Coordinate with on‑call SREs and the Monitoring and Incident Manager during…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).