×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Camden, Camden County, New Jersey, 08100, USA
Listing for: Sporttrade
Full Time position
Listed on 2026-08-29
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below

Our Story

Sporttrade is driven by a relentless pursuit to reduce costs and increase efficiency for market participants. Founded in 2018, Sporttrade is an aligned group of market technologists applying cutting edge market microstructure to provide unmatched quality of execution for customers.

Why Sporttrade?

Joining Sporttrade means partnering with a group of colleagues who share a deep respect for each other as we disrupt a multi-billion dollar emerging industry. While other companies will come and go, Sporttrade remains focused on providing best-in-class market structure guided by putting the customer first, always.

The Position

Sporttrade operates a regulated sports betting exchange that runs the way a financial exchange does. A successful Site Reliability Engineer at Sporttrade will be an operations-minded self-starter who treats the exchange like the production trading system it is: sessions open and close on schedule, every trade is captured and reported, and incidents are resolved quickly and documented thoroughly. In this role, you will own the day-to-day health of the exchange, spanning cloud infrastructure and on-premises datacenters, and you will automate away the toil that comes with operating a market that never wants to miss an open.

This role works closely with the Dev Ops, Technical Operations, and Engineering teams to keep a live, regulated marketplace fast, observable, and compliant.

Duties
  • Own the daily operation of the exchange trading lifecycle - market startup and shutdown, enabling and disabling trading, pre- and post-session sanity checks, and capture of settlement, clearing, and trade-reporting artifacts.
  • Participate in an on-call rotation for a live regulated marketplace; lead incident response, drive incidents to resolution, author postmortems, and turn one-off fixes into runbooks and automation
  • Operate and improve our observability stack (Datadog, Wazuh, Prometheus, Grafana) - dashboards, alert quality, SLOs, and reducing time-to-detection for market-impacting issues
  • Run and maintain hybrid infrastructure:
    Kubernetes clusters (GKE and EKS) with an Istio service mesh, AWS and GCP accounts, and exchange servers in geographically distributed on-premises datacenters
  • Automate infrastructure and operational procedures with Ansible, Terraform, and Jenkins pipelines, with secrets managed in Hashi Corp Vault
  • Support the data platform behind the exchange:
    PostgreSQL (Cloud SQL), Kafka (Confluent Cloud) change-data-capture and streaming pipelines, Redis, and backup/restore and disaster-recovery procedures — including proving that backups actually restore
  • Support market maker and partner connectivity (site-to-site VPNs and datacenter cross-connects), as well as conformance testing and onboarding support for partners joining the exchange.
  • Contribute to ongoing process improvement and the establishment of new policy and procedure for monitoring, incident management, change control, and exchange operations
Your portfolio
  • 5+ years of experience in a Site Reliability Engineering, Dev Ops, production engineering, or technical operations role supporting a 24/7 production system
  • Strong Linux fundamentals and scripting ability
  • Experience supporting and debugging Java applications in production - reading stack traces and thread dumps, working with JVM memory and garbage-collection behavior, and diagnosing service issues from logs and metrics
  • Solid working knowledge of TCP/IP networking — comfortable reasoning about connections, ports, routing, and firewalls to debug connectivity between exchange components, partners, and datacenters
  • Hands-on experience operating Kubernetes in production and managing infrastructure as code (Terraform, Ansible) with CI/CD pipelines (Jenkins or similar)
  • Experience with modern observability tooling (Datadog, Prometheus, Grafana, or equivalent) and a track record of being on-call for systems that matter
  • Working knowledge of SQL and relational databases (PostgreSQL preferred); experience with Kafka or other streaming platforms a plus
  • Self-starter who can deliver results with minimal guidance
  • Comfortable working independently and with a team
  • Excellent communication and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary