×
Register Here to Apply for Jobs or Post Jobs. X

Senior DevOps Engineer

Job in Hoboken, Hudson County, New Jersey, 07030, USA
Listing for: Akkadian Labs, LLC.
Full Time position
Listed on 2026-09-20
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below

IMPORTANT NOTE:
This role is only available to residents of the United States. We are unable to employ anyone who is not physically located in the US, nor are we sponsoring Visas at this time. Please be aware that job offers are contingent upon a background check which includes identity & address verification and a criminal background check.

Who We Are

Akkadian Labs is a Collaboration Lifecycle Automation Platform that services some of the largest global enterprises and government agencies. Our platform currently reduces manual work by up to 90% and costs by as much as 50%, while improving accuracy and governance across leading Unified Communications platforms.

This ability to innovate at scale defines who we are: big enough to compete at the highest level, yet agile enough to stay ahead of a rapidly evolving market. Our culture is people-first, fully remote, and rooted in grit, clarity, trust and respect, because when our people thrive, so do our customers.

Who You Are

You are a hands-on Senior Dev Ops Engineer with deep experience designing, building, and strengthening the infrastructure that keeps platforms secure, compliant, and uncomplicated. You bring substantial experience maintaining production infrastructure and improving operational reliability through disciplined engineering practice. You have a track record of helping to scale systems with a focus on dependability, observability, deployment confidence, and repeatable processes that support long-term growth.

You are energized by solving infrastructure challenges in complex enterprise environments and by building systems that reduce manual work, improve governance, and support operational accuracy at scale.

What You'll Do

This role supports Akkadian’s purpose to uncomplicate work by removing friction between people and their purpose — helping enterprise customers govern, automate, and continuously control complex collaboration environments across cloud, hybrid, and on-premises systems. You will design, implement, and maintain scalable and secure infrastructure and Dev Ops processes at Akkadian Labs. You will work closely with development, QA, and product teams to enable reliable deployments, automate workflows, and improve system observability across Rocky OS-based, AWS-hosted, and on-premises solutions.

This is a hands-on technical role focused on system-level execution, continuous improvement, and operational excellence within the Dev Ops function.

Key Responsibilities Infrastructure and Environment Management
  • Deploy and maintain scalable infrastructure in AWS and hybrid cloud environments.
  • Manage infrastructure-as-code (IaC) using Terraform, Cloud Formation, or similar tools.
  • Maintain Linux-based environments.
  • Design and implement containerization using Docker and orchestration via Kubernetes.
AI and Agent Infrastructure Implementation & Support
  • Design, deploy and manage AI agent workloads, including provisioning compute instances and managing resource scaling for inference-heavy tasks.
  • Build and maintain model deployment pipelines, including versioning, testing, and rollback of AI models in production environments.
  • Monitor AI API consumption and infrastructure costs, implementing alerting and controls to prevent runaway usage and support budget visibility.
  • Collaborate with the engineering team to implement infrastructure-level security guardrails for AI systems, including access controls and data isolation for model inputs and outputs.
Observability and Reliability
  • Manage monitoring and observability efforts using tools such as Prometheus, Grafana, and the ELK stack.
  • Troubleshoot system issues and contribute to incident response and root cause analysis.
  • Develop and execute strategies for improving system reliability, performance, and…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary