Director of Operations, Infrastructure & SRE
Listed on 2026-08-03
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
About Us
Mintegral is a leading programmatic and interactive mobile advertising platform. Focused on the APAC region and radiating out globally. Powered by advanced AI technology, we provide global advertisers and developers with innovative, comprehensive experiences. With our efficient mobile marketing and monetization solutions, we help our clients exceed their marketing goals.
Core ObjectivesBuild Mintegral’s automated operations framework to improve deployment, monitoring, alerting, incident response, capacity planning, and resource management efficiency.
Optimize server, cloud, bandwidth, storage, and compute costs, improve resource utilization, and reduce cost per unit of traffic processed.
Improve the overall operating efficiency of global advertising systems, including request processing capacity, system latency, availability, elastic scaling, and failure recovery.
Support Mintegral’s global business growth by ensuring infrastructure and operations capabilities can efficiently support ad traffic, customer needs, and commercialization goals.
Duties and Responsibilities- Lead the planning, development, and continuous optimization of Mintegral’s global operations, infrastructure, and SRE systems to support large-scale ad requests, real-time bidding, and data processing systems.
- Build and improve automated operations platforms and processes, including automated deployment, monitoring and alerting, capacity planning, incident management, resource scheduling, configuration management, cost analysis, and performance optimization.
- Drive server and cloud infrastructure cost optimization, including resource pooling, elastic scaling, capacity governance, cost attribution, procurement strategy, and hybrid-cloud / multi-cloud architecture, with the goal of continuously reducing infrastructure costs.
- Improve core system efficiency and reliability by optimizing the ad request pipeline, infrastructure performance, network latency, service availability, disaster recovery capabilities, and global node deployment strategy.
- Establish a comprehensive observability framework, including logs, metrics, distributed tracing, real-time monitoring, SLA/SLO management, incident reviews, and risk early-warning mechanisms.
- Work closely with engineering, algorithms, product, data, sales, and commercialization teams to ensure operations and infrastructure capabilities support business growth, customer delivery, and global market expansion.
- Drive operations automation and intelligence by introducing or building AIOps, automated root cause analysis, anomaly detection, capacity forecasting, and intelligent resource scheduling capabilities.
- Establish core metrics for cost, efficiency, and reliability, and regularly report infrastructure ROI, optimization progress, and business impact to the CEO and senior leadership team.
- Depending on the level, lead or guide operations, SRE, infrastructure, or platform engineering teams, including technical management, talent development, and organizational planning.
- 5+ years of experience in operations, SRE, infrastructure, platform engineering, or technical operations, preferably in large-scale internet platforms, ad tech platforms, cloud platforms, or distributed systems.
- Proven success building automated operations systems, SRE frameworks, infrastructure platforms, or cost optimization mechanisms, with measurable impact on reliability, efficiency, or cost.
- Strong knowledge of large-scale distributed systems, high-concurrency services, low-latency systems, global deployment, containerization, Kubernetes, cloud services, CDN, network optimization, and capacity planning.
- Experience optimizing server costs, with familiarity in cloud cost governance, resource utilization optimization, Fin Ops, capacity management, elastic scaling, procurement strategy, or hybrid-cloud / multi-cloud architecture.
- Strong familiarity with monitoring and alerting, log analysis, distributed tracing, incident response, SLA/SLO management, disaster recovery, high-availability architecture, and performance tuning.
- Strong data analysis skills and business awareness, with the ability to identify optimization…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).