×
Register Here to Apply for Jobs or Post Jobs. X

Customer Support Engineer

Job in Cambridge, Middlesex County, Massachusetts, 02140, USA
Listing for: Blitzy
Full Time position
Listed on 2026-07-14
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, IT Support, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 100000 - 140000 USD Yearly USD 100000.00 140000.00 YEAR
Job Description & How to Apply Below

About Blitzy

Blitzy is a Cambridge, MA based AI software development platform on a mission to revolutionize the software development life cycle by autonomously building custom software to unlock the next industrial revolution. We're transforming how enterprises build software, turning enterprise requirements into production-ready code with an agentic software development platform that can autonomously execute 80% of the quantum of software development work.

We're backed by multiple tier-1 investors, and have proven success as founders of previous start-ups.

Location and Compensation

Location:

Cambridge, MA (On-site)

Compensation: $100,000 - $140,000 salary + equity

The Role

The role is to support our clients and ensure a stable environment across the full lifecycle: installation, ongoing upgrades, and day-to-day operation. The L2 Support Engineer works alongside L1 to triage and resolve issues, and escalates unresolved defects to engineering. It operates across Kubernetes, Docker, and the major cloud providers.

What Success Looks Like
  • Customers' issues are resolved faster and escalated cleaner.
  • Recurring problems turn into runbooks, dashboards, and alerts, not repeat tickets.
  • Engineering trusts your escalations because they come with proof, not guesses.
  • Customers trust your communication because it's clear, honest, and on time.
Areas of Ownership
  • Deploy and install the platform into customer environments, and troubleshoot installation issues.
  • Support ongoing upgrades and day-to-day operation, keeping customer environments stable.
  • Work alongside L1 to triage and resolve customer-reported issues, driving them to resolution or escalation.
  • Diagnose failures across the stack: compute, networking, storage, and the services running on it.
  • Reproduce issues safely against live (often multi-tenant) environments using read-only diagnostics first.
  • Build and maintain dashboards, monitors, and runbooks so recurring issues get faster to fix: or stop recurring.
  • Write up clear, evidence-backed escalations and post-incident notes.
  • Communicate status and resolution to customers clearly and on time.
Required Experience
  • Distributed-systems debugging. Reason about a request crossing multiple services, queues, and network hops, and isolate which hop failed. You debug by forming a hypothesis and confirming it with evidence (logs, pod state, queue depth, DB rows), not by guessing.
  • Kubernetes & Docker.
  • Major cloud providers: GCP, AWS, and Azure. Hands-on with at least one deeply and able to work across the others: managed Kubernetes (GKE/AKS), cloud logging, IAM/auth basics, and cloud disk/storage behavior.
  • Strong monitoring & observability practice. Fluent with an APM/observability stack (Datadog or equivalent): log queries, correlating across services by request/trace IDs, reading traces, and building dashboards and alerts. You reach for the data before theorizing.
Additional Skills & Experience
  • Python and Redis literacy.
  • Basic message queueing. Command transport runs over a message queue (Redis/rq). Comfort inspecting queue depth, backlogs, and stuck/failed jobs; concepts transfer from any broker.
  • Networking & Web Sockets. Many of our hardest issues are connection problems:
    Web Socket/Socket.

    IO drops, NAT/idle/LB timeouts, half-open sockets, DNS-vs-routing, TLS. Tell a transport fault from an application fault.
  • SQL / PostgreSQL. Query operational tables to confirm what the system recorded.
  • Source-control platforms. Git Hub (incl. Git Hub Enterprise Server), Azure Dev Ops, and/or Git Lab, clone/push/pull, access tokens, app credentials, and their failure modes.
  • CI/CD, Helm & deploy integrity. Many “sudden regressions” are a bad or partial deploy: check what version is actually running before chasing architecture theories. Helm and container deploy pipelines expected. ArgoCD is a plus.
  • Secrets management. Comfort handling secrets, credentials, and certificates safely, ideally with Vault (strongly preferred).
  • Linux and Windows. Workloads run on both; comfort triaging on each OS (process inspection, file system, basic networking).
  • Methodical, evidence-first temperament. Hold several candidate causes at once, run the cheapest disconfirming check…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary