×
Register Here to Apply for Jobs or Post Jobs. X

Observability Operations Engineer

Job in Phoenix, Maricopa County, Arizona, 85003, USA
Listing for: BC Forward
Full Time position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 52 - 60 USD Hourly USD 52.00 60.00 HOUR
Job Description & How to Apply Below
Job Title:

Observability Operations Engineer

Location:

(City, State) Duration:
Temp - 12 months Pay Range: $52/hr $60/hr (W2) Job  About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity.

Job Description We are seeking a Senior Observability Operations Engineer to join our dynamic team. The ideal candidate will have strong experience in Dynatrace, Splunk, Open Search/Elasticsearch, Kubernetes, Linux, and cloud-native observability and a proven ability to automate operations with AI/ML and improve incident response and platform reliability. Responsibilities:
Administer and optimize Dynatrace, Splunk, and Open Search/Elasticsearch for enterprise observability. Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions. Manage large Open Search/Elasticsearch clusters, including indexing, performance tuning, shard optimization, backups, and capacity planning. Configure Dynatrace One Agent, Active Gate, Synthetic Monitoring, RUM, DEM, Davis AI, and APM. Administer Splunk Enterprise, Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and ITSI.

Develop dashboards, alerts, reports, and executive operational metrics. Support Linux infrastructure and Kubernetes environments, including Docker, Open Shift, or Rancher. Implement observability best practices using Open Telemetry for traces, metrics, logs, and events. Perform root cause analysis for production incidents using observability platforms. Collaborate with Platform Engineering, SRE, Dev Ops, Infrastructure, and Application teams. Automate operational tasks using Python, Shell, REST APIs, Terraform, or Ansible.

Role mix is approximately 20% automation and 80% operations. Participate in incident, problem, change, and release management processes. Drive upgrades, patching, security compliance, and operational governance. Improve reliability through automation, self-healing, and AI-assisted operations. Required Skills &

Qualifications:

Dynatrace, Splunk Enterprise, and Open Search/Elasticsearch administration. Grafana, Prometheus, Kibana, Jaeger, and Open Telemetry. Kafka preferred. Linux administration, Kubernetes, Docker, Open Shift or Rancher, and core networking. AWS, Azure, or GCP; CI/CD;
Git;
Terraform;
Ansible; REST APIs. Python and Bash/Shell. Power Shell preferred. Experience supporting enterprise production environments with strong troubleshooting and analytical skills. Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience. 6-10+ years IT or observability operations and 4+ years administering Dynatrace, Splunk, Open Search, or Elasticsearch. Preferred

Skills:

AI and automation experience, including Dynatrace Davis AI and AIOps practices. Generative AI tools such as ChatGPT, Git Hub Copilot, Amazon Q, or Microsoft Copilot to improve operational efficiency. Machine learning concepts for predictive monitoring and intelligent alerting. RAG, vector databases, embeddings, and AI-powered search are a plus. Python with AI frameworks such as Lang Chain, Lang Graph, or OpenAI APIs.

Experience integrating AI with observability via APIs. Kubernetes AI/ML, Grafana, and Sahara familiarity. Work Schedule & Expectations:
Onsite three days per week. Standard shifts are 9:00 AM-6:00 PM or 10:00 AM-6:30 PM for coverage. Willingness to work one weekend day every two to three weeks. Preferred

Certifications:

Dynatrace Associate or Professional. Splunk Enterprise Certified Administrator. Elastic Certified Engineer. Kubernetes CKA or CKAD. AWS, Azure, or GCP certification. ITIL Foundation. AI/ML or Generative AI certification preferred.

Soft Skills:

Ownership, accountability, and strong problem-solving. Independent work with minimal supervision and effective collaboration.…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary