More jobs:
AI Ops Engineer
Job in
Bryan, Brazos County, Texas, 77808, USA
Listed on 2026-07-04
Listing for:
TechDigital Group
Full Time
position Listed on 2026-07-04
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support, Systems Administrator
Job Description & How to Apply Below
Experience Requirements
- 5+ years in IT operations or L1 support roles.
- Exposure to AIOps environments or automated monitoring solutions is a plus.
- Bachelor's or master's degree in computer science, Engineering, or a related field.
Splunk, Power Shell, or Python, Logs Monitoring, Confluence and Share Point
Skill Requirements- Hands‑on experience with IT monitoring tools (e.g., Nagios, Zabbix, Prometheus, Splunk, or similar).
- Understanding of scripting (Power Shell, Python, or Shell) for basic automation tasks.
- Understanding of AIOps concepts and automation frameworks.
- Proficiency in Confluence and SharePoint for status updates and documentation.
- Ability to interpret logs and detect anomalies proactively.
- Familiarity with ITIL processes for incident, problem, and change management.
- Experience using ticketing systems (e.g., Service Now, Jira, Remedy).
- Skilled in creating and updating runbooks and SOPs.
- Ability to follow documented procedures accurately.
- Strong attention to detail for maintaining health check reports and incident updates.
- Analytical thinking for quick problem identification and escalation.
- Excellent communication and documentation skills.
- Proactive mindset with a passion for reliability and automation.
- Strong problem‑solving and debugging skills.
- ITIL Foundation Certification.
- Experience with anomaly detection, time‑series forecasting, and log analysis.
- Basic certifications in monitoring tools or cloud platforms (AWS, Azure).
- Proactive Monitoring of alerts and detect anomalies from logs.
- Perform daily health checks until full automation and application monitoring are implemented.
- Follow status checks as per existing runbooks.
- Create and update runbooks as needed to reflect current processes.
- Update system health status every 2 hours during the shift in Confluence or SharePoint.
- Acknowledge incidents promptly and route them to the correct team.
- Update incident status every 4 hours for P1/P2 tickets.
- Communicate with users and provide timely updates on their requests.
- Ensure timely acknowledgment, follow‑up, and closure of incidents within SLA.
- Complete service tasks on time as per SLA to release queues quickly.
- Work strictly as per SOPs documented by the team.
- Familiarity with incident management processes and ITIL principles.
- Ability to follow documented procedures and create/update runbooks.
- Strong communication and coordination skills.
- Understanding of Confluence, SharePoint, and ticketing systems.
- Implement best practices in ML operations and productionization.
- Ensure compliance with enterprise data security, governance, and regulatory requirements.
- Collaborate with data engineers, analysts, Dev Ops/SRE teams and business teams to ensure reliability and security.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×