×
Register Here to Apply for Jobs or Post Jobs. X

Software Engineer, MTIA SW Tooling Engineer

Job in Menlo Park, San Mateo County, California, 94029, USA
Listing for: Meta
Full Time position
Listed on 2026-08-19
Job specializations:
  • Software Development
    Software Engineer, AI Engineer (Applied/Software), DevOps
Salary/Wage Range or Industry Benchmark: 184000 - 257000 USD Yearly USD 184000.00 257000.00 YEAR
Job Description & How to Apply Below

Meta is seeking a Software Engineer to join the MTIA (Meta Training & Inference Accelerator) Software Tooling team, which develops and maintains the tooling ecosystem for Meta's in-house AI accelerator ASICs. The Tooling team provides debugging, profiling, memory analysis, and monitoring capabilities for the whole MTIA Ecosystem, advancing ML accelerator tooling by leveraging Meta's full-stack ownership from silicon specs to fleet observability.

In this role, you will be a senior technical contributor responsible for designing and building developer tools that help engineers debug, profile, measure, and monitor AI workloads running on MTIA hardware  will work at the intersection of compilers, runtime, hardware, and ML frameworks, collaborating with cross-functional partners to deliver a high-quality developer experience for Meta's custom AI accelerators.

Software Engineer, MTIA SW Tooling Engineer Responsibilities:
  • Own the technical vision and roadmap for key areas of MTIA's developer tooling ecosystem, focusing on debugging, workload error analysis, and product / fleet reliability
  • Design and develop debugging tools - including graph-mode debugging, kernel-level diagnostics, and multi-rank fault isolation
  • Contribute to overall MTIA SW tooling infrastructure - enabling profiling, performance debugging, memory sanitization, and reliability analysis for MTIA training and inference workloads
  • Collaborate closely with MTIA compiler, runtime, kernel, and hardware teams to integrate tooling hooks throughout the MTIA software stack
  • Drive AI-native tooling approaches by leveraging automation and LLM-guided diagnostics to improve developer productivity and reduce time-to-root-cause
  • Partner with internal product teams across advertising, recommendations, and generative AI to understand developer pain points and prioritize tooling investments
  • Advise on tooling best practices, debugging methodologies, and systems-level analysis for accelerator software
  • communicate architectural decisions clearly through design documents and cross-team reviews
Minimum Qualifications:
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 6+ years of experience in systems software engineering, performance engineering, developer tooling, or a closely related field
  • Experience building debugging, profiling, or diagnostic tools for complex software/hardware systems
  • Proficiency in C++ and Python, including low-level systems programming and scripting for tool automation
  • Experience working across multiple layers of a system stack (compiler, runtime, OS/driver, hardware)
  • Experience leading the technical design and delivery of tooling or infrastructure projects from inception through production deployment
  • Experience using data-driven methods and experimentation to evaluate and validate tooling effectiveness and systems performance improvements
Preferred Qualifications:
  • Familiarity with ML framework internals (PyTorch graph execution, torch.compile, operator dispatch) and AI compiler stacks (MLIR, LLVM, TVM, Triton)
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Experience with accelerator ecosystems (GPU/CUDA, TPU, custom ASICs) including performance profiling, memory analysis, and runtime debugging using their tool chains (cuda-gdb, nsight-compute, nsight-systems, cuda-memcheck)
  • Demonstrated cross-stack debugging ability, including Linux kernel and driver-level debugging, with capacity to trace issues across application, OS, and hardware boundaries
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Experience with distributed systems debugging or profiling (multi-device, multi-node/multi-rank)
  • Experience with Linux debugging and profiling infrastructure (gdb, perf, eBPF, ftrace, coredump analysis, hardware performance counters) and familiarity with binary formats and debugging metadata (ELF/DWARF)
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary