Senior Site Reliability Engineer, Platform Responsibility - USDS
Listed on 2026-07-26
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Cybersecurity
Responsibilities
Team Intro:
The Platform Responsibility engineering team is fast growing and responsible for building machine learning models and systems to identify and defend internet abuse and fraud on our platform. Our mission is to protect billions of users and publishers across the globe every day. We embrace the state-of-the-art machine learning technologies and scale them to detect and improve trust and safety system using the tremendous amount of data generated on the platform.
With the continuous efforts from our team, Tik Tok USDS is able to provide the best user experience and bring joy to everyone in the world.
- Manage day-to-day operations of data service, realtime/batch data pipelines, such as SLA/SLO/SLI management, system deployment, performance tuning and troubleshooting
- Design and deploy AI Agents and LLM-powered automation to streamline incident response, root cause analysis, and proactive system monitoring
- Create tools and automation to improve system administration and operational efficiency, leveraging AI-assisted development tools to accelerate delivery and code quality
- Participate in regular on-call rotations as part of a team that provides 24 hour coverage across multiple shifts
- Engage in and improve the whole lifecycle of services from inception and design, development, capacity planning, and launch reviews, to deployment, operation, and refinement
- Practice sustainable user support, incident response, and post mortem
- Drive the long-term roadmap for system administration tools. Build internal platforms that leverage AI-assisted development to eliminate toil and improve engineering velocity.
- Serve as a primary Incident Commander for high-severity issues, leading cross-functional teams and ensuring technical resolution aligns with business priorities.
- Partner with Product and Development teams from the design phase to ensure observability, scalability, and disaster recovery are core components of every new feature.
- Foster a culture of sustainable operations by mentoring junior engineers and evangelizing SRE principles throughout the organization.
Minimum Qualifications:
- Bachelor’s degree or above in Computer Science or a related technical discipline.
- At least 5 years of professional experience in SRE, Dev Ops, or Infrastructure Engineering.
- Proven experience integrating AI/LLM APIs into internal workflows, specifically for log analysis, alert contextualization, or diagnostic assistance.
- Deep understanding of Unix/Linux system internals, networking fundamentals, and distributed systems architecture.
- Expertise in designing and scaling observability stacks using tools such as Prometheus, Grafana, or Data Dog.
- Demonstrated ability to troubleshoot complex, non-obvious production issues across the entire stack.
Preferred Qualifications:
- Experience building agentic workflows and orchestration frameworks to assist in incident triaging and runbook matching.
- Mastery of AI-assisted development tools to accelerate infrastructure-as-code delivery and automate documentation.
- Deep technical proficiency in container orchestration via Kubernetes and managing big data technologies such as Kafka.
Tik Tok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order signed by President Trump on September 25, 2025. Our foundation is a comprehensive data privacy and cybersecurity program we operate under defined safeguards to protect national security and secure U.S. user data, apps and the algorithm.
We safeguard the U.S. content ecosystem, holding decision‑making authority for trust and safety policies and moderation. USDS Joint Venture helps ensure Americans can continue to express their creativity, discover new hobbies and interests, and build thriving communities and businesses on a global scale.
On‑site presence across teams allows the company to operate with greater speed, alignment, and agility — especially in areas like real‑time decision‑making, team development, and integrated execution. As such,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).