×
Register Here to Apply for Jobs or Post Jobs. X

Manager, Product Engineering

Job in Fort Lauderdale, Broward County, Florida, 33316, USA
Listing for: Grand Turk Cruise Center
Full Time position
Listed on 2026-09-06
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below

Manager, Web and Mobile Site Reliability Engineering

One of the best-known names in cruising, Princess is the world's leading international premium cruise line and tour company, carrying millions of guests each year to hundreds of destinations around the globe. We give our guests the Medallion Class experience others simply can't. The Love Boat promises something for everyone.

The Manager, Web and Mobile Site Reliability Engineering (SRE) leads the engineering team responsible for ensuring maximum uptime, high availability, performance, and resilience for enterprise web applications, mobile app backends, and public API endpoints. This role defines reliability standards, oversees 24/7 incident response, manages edge infrastructure and bot mitigation, and drives automated deployment and observability pipelines.

Here's a summary of what Princess is looking for in a Manager, Web and Mobile Site Reliability Engineering. Is this you?

Responsibilities:

  • Team Leadership & SRE Operations:
    Lead and develop a high-performing team of SRE and Dev Ops engineers supporting 24/7 high-volume web and mobile systems. Manage on-call rotations, incident command protocols, and operational readiness.
  • Reliability & Observability Governance:
    Establish Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. Architect end-to-end monitoring, tracing, and alerting strategies using tools like Datadog, Dynatrace, or Grafana.
  • Incident Management & Remediation:
    Lead major incident response efforts, drive blameless post-mortems, and collaborate with engineering teams to prioritize root-cause fixes and architectural resiliency improvements.
  • Traffic, Edge & Security Management:
    Partner with IT Security (PCL IT Security) and CDN providers (Akamai) to implement bot mitigation strategies, DDoS defense, WAF rules, and edge caching for key APIs and digital endpoints.
  • Administrative:
    Perform all other administrative and organizational duties as required (time keeping, training, travel, collaboration and correspondence, etc.)

Knowledge &

Skills:

  • Scope:
    Direct management of SRE and Dev Ops engineers. Operational oversight for consumer-facing web platforms, mobile backend APIs, edge routing networks, and cloud deployment pipelines.
  • Problem Solving:
    Rapidly diagnoses and mitigates complex system outages, performance bottlenecks, traffic anomalies, bot campaigns, and infrastructure failures in high-volume production environments. Resolves highly complex, enterprise-scale operational challenges that impact guest operations, maritime services, revenue-generating systems, regulatory requirements, and technology service availability. Anticipates emerging operational risks, evaluates competing business priorities, establishes governance frameworks, and makes decisions where significant operational, financial, service, and reputational consequences may exist.

    Develops innovative solutions to improve enterprise resilience, scalability, and operational effectiveness.
  • Impact:
    Directly ensures continuous operational availability, system security, optimal site performance, and guest trust across web and mobile touchpoints.
  • Leadership:
    The role requires strong leadership skills. Requires strong incident command leadership, strategic operational decision-making, calm under pressure, and collaborative mentorship.
  • Knowledge:
    In-depth understanding of Site Reliability Engineering practices, cloud platforms (AWS/Azure), containerization (Kubernetes, Docker), Akamai/CDN edge routing, bot detection, and CI/CD pipelines (Git Lab).
  • Skills:

    Production incident management, automated infrastructure management (Terraform), performance tuning, distributed tracing, metrics-driven SLI/SLO establishment.
  • Abilities:
    Ability to lead teams during critical production outages, drive cross-functional engineering accountability for reliability, and automate operational workflows.

Essential/

Minimum Qualifications:

  • Bachelor's degree in Computer Science, Computer Engineering, System Administration, or equivalent experience.
  • 6+ years in Site Reliability Engineering, Dev Ops, or Infrastructure Engineering.
  • 2+ years of leadership or direct engineering management…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary