×
Register Here to Apply for Jobs or Post Jobs. X

AI Ops Engineer

Job in Bryan, Brazos County, Texas, 77808, USA
Listing for: TechDigital Group
Full Time position
Listed on 2026-07-04
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support, Systems Administrator
Salary/Wage Range or Industry Benchmark: 70000 - 90000 USD Yearly USD 70000.00 90000.00 YEAR
Job Description & How to Apply Below

Experience Requirements

  • 5+ years in IT operations or L1 support roles.
  • Exposure to AIOps environments or automated monitoring solutions is a plus.
Qualifications
  • Bachelor's or master's degree in computer science, Engineering, or a related field.
Key Skills

Splunk, Power Shell, or Python, Logs Monitoring, Confluence and Share Point

Skill Requirements
  • Hands‑on experience with IT monitoring tools (e.g., Nagios, Zabbix, Prometheus, Splunk, or similar).
  • Understanding of scripting (Power Shell, Python, or Shell) for basic automation tasks.
  • Understanding of AIOps concepts and automation frameworks.
  • Proficiency in Confluence and SharePoint for status updates and documentation.
  • Ability to interpret logs and detect anomalies proactively.
  • Familiarity with ITIL processes for incident, problem, and change management.
  • Experience using ticketing systems (e.g., Service Now, Jira, Remedy).
  • Skilled in creating and updating runbooks and SOPs.
  • Ability to follow documented procedures accurately.
  • Strong attention to detail for maintaining health check reports and incident updates.
  • Analytical thinking for quick problem identification and escalation.
  • Excellent communication and documentation skills.
  • Proactive mindset with a passion for reliability and automation.
  • Strong problem‑solving and debugging skills.
Preferred
  • ITIL Foundation Certification.
  • Experience with anomaly detection, time‑series forecasting, and log analysis.
  • Basic certifications in monitoring tools or cloud platforms (AWS, Azure).
Key Responsibilities
  • Proactive Monitoring of alerts and detect anomalies from logs.
  • Perform daily health checks until full automation and application monitoring are implemented.
  • Follow status checks as per existing runbooks.
  • Create and update runbooks as needed to reflect current processes.
  • Update system health status every 2 hours during the shift in Confluence or SharePoint.
  • Acknowledge incidents promptly and route them to the correct team.
  • Update incident status every 4 hours for P1/P2 tickets.
  • Communicate with users and provide timely updates on their requests.
  • Ensure timely acknowledgment, follow‑up, and closure of incidents within SLA.
  • Complete service tasks on time as per SLA to release queues quickly.
  • Work strictly as per SOPs documented by the team.
  • Familiarity with incident management processes and ITIL principles.
  • Ability to follow documented procedures and create/update runbooks.
  • Strong communication and coordination skills.
  • Understanding of Confluence, SharePoint, and ticketing systems.
  • Implement best practices in ML operations and productionization.
  • Ensure compliance with enterprise data security, governance, and regulatory requirements.
  • Collaborate with data engineers, analysts, Dev Ops/SRE teams and business teams to ensure reliability and security.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary