Site Reliability Engineer
Position Summary:
We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer, you will play a crucial role in enhancing the reliability, performance, and scalability of our systems and services. You will be a part of a global “commando” team of highly skilled SREs, driving best practices and innovations for optimal system operations, while protecting critical companies’ systems in real time.
Whyjoin Tipalti?
Tipalti is the AI-powered platform for finance automation, elevating how finance teams operate in the global economy. We empower our customers to scale faster and smarter by removing the complexities of doing global business and accelerating their finance operations efficiency. Our platform provides a comprehensive suite of finance automation solutions designed for mid-market businesses across accounts payable, global payouts, procurement, employee expenses, corporate cards, supplier management, tax compliance, and treasury.
Tipalti partners with leading financial institutions such as Citi, Wells Fargo, J.P. Morgan, and Visa, enabling over 5,000 global companies to efficiently and securely pay millions of suppliers and payees across 200+ countries and territories, in 120 currencies.
At Tipalti, we pride ourselves on our collaborative culture, the quality of our product and the capabilities of our people. Tipaltians are passionate about the work they do, and keen to get the job done. Tipalti offers competitive benefits, a flexible workplace, career coaching, and an environment where diverse individuals can thrive and make an impact. Our culture ensures everyone checks their egos at the door and stands ready to reach for success together.
Founded in Israel in 2010, Tipalti is a global business headquartered in the San Francisco Bay Area (Foster City) with offices in Tel Aviv, Plano, Toronto, Vancouver, London, Amsterdam, Tbilisi and Medellin.
In this role, you will be responsible for:- Driving incident response and post-mortem processes, fostering a culture of continuous improvement.
- Designing, building and improving internal tools and automation software to make maintaining production services easier and safer.
- Lead reliability-focused practices such as SLO (Service Level Objective) design and implementation, Failure Analysis, Load and Capacity Planning, Service Reviews, Architecture Designs, Incident Postmortems, and others.
- Participate in the on-call rotation, providing expertise and support during critical system incidents and ensuring timely resolution.
Required:
- Minimum 4+ years of Software Engineering experience in an object-oriented or modern language (preferably .NET or Type Script), with a proven track record of strong learning agility
- Experience working on large-scale, high-traffic platform for an end-product company
- Demonstrate learning agility and willingness to learn new technologies
- Hands-on experience or familiarity with production deployment and end-to-end lifecycle ownership of applications
- Knowledge of architecture and application design experience.
- Solid troubleshooting and debugging skills.
- Excellent verbal and written communication skills in English.
- Basic understanding of some of the lower levels of software frameworks and networking.
- Distributed monitoring experience with logging, metrics and tracing using Open Telemetry and Prometheus.
- Additional scripting languages: bash, powershell, python
- Previous experience working as SRE
- Experience with working in a cloud driven environment (AWS, GCP, Azure)
- Hybrid working model that requires you to be in office Monday, Tuesday, and Thursday
- Competitive salary and equity
- Matching RRSP
- Extended benefits,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).