Infrastructure Engineer
Listed on 2026-08-11
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AWS
About Lineate
Lineate is a US-based international software development company with over two decades of experience.
From Intelligent Document Processing(IDP) and Agentic RAG systems to scalable cloud architectures, we turn complex ideas into real, measurable results.
We deliver AI-driven custom solutions for Fin Tech, Health Tech, AdTech, and beyond, empowering businesses to grow smarter, faster, and more efficiently.
Our expertise falls into three main categories:
- Building Custom AI Solutions:
Deploying high-impact, AI-enabled technology utilizing IDP, Agentic RAG. - Cloud and Data Infrastructure:
Optimizing business operations with our data management and cloud computing solutions. - Team Augmentation:
Providing specialized experts in Fin Tech, AdTech, and Health Tech to integrate seamlessly and accelerate project timelines. - Our goal is not just to build technology, but to build the future operating model for our clients.
The Infrastructure Engineer is responsible for the reliability, security, and continuous improvement of Lineate and customer infrastructure across cloud and on-prem environments. The role combines hands-on SRE/Dev Ops work - automation, observability, incident response - with support engineering for production systems. The engineer is expected to operate independently on most tasks, and drive infrastructure and tooling improvements, including the use of AI agents to automate routine operational work.
Accountablefor
- Reliability, availability, and performance of Lineate and customer production systems in accordance with agreed SLOs and support SLAs
- Design, automation, and continuous improvement of cloud and on-prem infrastructure
- Timely and high-quality response to incidents and escalated support requests
- Security, compliance, and cost-efficiency of the infrastructure under the team's ownership
- Day-to-day operations of the infrastructure owned by the team
- Participate in a 24/7 shift schedule
- Monitor systems and respond to automated alerts in line with Lineate support SLAs
- Diagnose and resolve system problems
- Operate Kubernetes workloads and Databricks resources by following team playbooks and runbooks
- Perform routine system administration on Linux and Windows servers
- Respond promptly and professionally to customer inquiries via Slack, email, or Zendesk
- Document, categorize, and prioritize incoming support requests
- Communicate the status of open requests clearly and regularly to customers
- Identify and implement automation opportunities, including AI-agent-driven automation of routine operational tasks (ticket triage, runbook execution, log analysis, knowledge-base maintenance)
- Drive incident resolution by coordinating triage, implementing mitigation measures, conducting root cause analysis, and leading post-incident reviews
- Maintain and improve runbooks, playbooks, FAQs, and internal knowledge bases
- Implement and improve observability: metrics, logs, traces, dashboards, alerting, and SLO tracking
- Contribute to CI/CD pipelines and basic infrastructure automation
- Follow basic information security practices (IAM, secrets handling, least privilege)
- Participate in the improvement of support and infrastructure processes
- Minimum 3 years of hands-on Linux system administration experience (required)
- Minimum 2 years of hands-on experience with AWS in a production environment (required)
- Working familiarity with Kubernetes — able to operate workloads, read logs, and perform standard actions by following playbooks (deep cluster-design expertise not required)
- Working familiarity with AWS Databricks — able to run jobs, manage clusters, and troubleshoot using existing playbooks (deep platform expertise not required)
- Practical experience using AI agents / LLM-based tooling (e.g., Claude, ChatGPT, MCP-based agents, Cursor, Copilot) to automate operational, scripting, or support tasks
- At least 1 year of Windows system administration experience
- Practical scripting skills in at least one of Python, Bash, Go (1+ years of practical use)
- Ability to maintain existing pipelines in CI/CD tools such as Jenkins, Git Hub Actions, Git Lab CI, or equivalent
- Proficiency with Git and standard version-control workflows
- Hands-on experience with Docker and container workflows
- Experience managing and maintaining production IT infrastructure
- Comfortable working with a broad range of open-source technologies and tools
- Infrastructure as Code:
Terraform (preferred), Ansible, or Cloud Formation - Observability stack:
Prometheus, Grafana, ELK/Loki, Datadog, Cloud Watch, or equivalent—including building dashboards and alerts - Networking fundamentals: TCP/IP, DNS, HTTP(S), TLS, VPN, load balancers, VPC design ,basic firewall and routing
- Cloud security: IAM, KMS, secrets management (AWS Secrets Manager, Hashi Corp Vault), least-privilege design, audit logging
- Database operations:
PostgreSQL, MySQL, or Redis — backups, restores, basic tuning, monitoring - Incident management practices: on-call rotations, runbooks,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).