AI Engineer/DevOps Engineer
Listed on 2026-08-30
-
IT/Tech
SRE/Site Reliability
Key Responsibilities:
Platform & Application Monitoring:
- Monitor production applications and Azure-based services to ensure availability, reliability, and performance.
- Review observability dashboards and alerts to identify abnormal conditions and proactively prevent incidents.
- Latency and response times
- User interface accessibility and availability
- API and endpoint health
- Validate service-level metrics and
- Monitor and maintain Azure infrastructure components including:
- Perform routine operational activities and platform administration.
- Identify infrastructure bottlenecks and recommend remediation actions. [
- Monitor Kubernetes clusters and containerized workloads.
- Troubleshoot pod failures, resource utilization issues, and deployment problems.
- Perform operational recovery activities including:
- Restarting services
- Support high availability and platform stability initiatives. [
- Monitor and troubleshoot Azure Dev Ops CI/CD pipelines.
- Investigate deployment failures and integration issues.
- Support Dev Ops service connections, agent pools, pipeline execution, and release processes.
- Ensure successful movement of code and configurations across environments.
- Validate source file arrivals and data ingestion processes.
- Troubleshoot failed jobs, ETL processes, and data movement issues.
- Perform data validation and staging activities.
- Ensure data pipelines meet operational and business requirements.
- Respond to incidents and service disruptions in accordance with established SLAs.
- Perform incident triage, analysis, and resolution coordination.
- Create detailed Root Cause Analysis (RCA) documentation for production issues.
- Develop and track Corrective Action Plans (CAPs).
- Participate in problem management and continuous improvement activities.
- Support change validation and deployment activities.
Identify automation opportunities to reduce manual effort.
Leverage observability insights to improve system reliability and operational efficiency.
Contribute to monitoring enhancements, alert optimization, and operational runbooks.
Maintain support documentation, SOPs, and knowledge articles.
The pay range that the employer in good faith reasonably expects to pay for this position is $47.15/hour - $73.68/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.
Tundra Technical Solutions is among North America’s leading providers of Staffing and Consulting Services. Our success and our clients’ success are built on a foundation of service excellence. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Unincorporated LA County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: client provided property, including hardware (both of which may include data) entrusted to you from theft, loss or damage;
return all portable client computer hardware in your possession (including the data contained therein) upon completion of the assignment, and; maintain the confidentiality of client proprietary, confidential, or non-public information. In addition, job duties require access to secure and protected client information technology systems and related data security obligations.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).