×
Register Here to Apply for Jobs or Post Jobs. X

Chip Profiling Engineer - Member of Technical Staff

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Infinity Artificial Intelligence Institute
Full Time position
Listed on 2026-07-22
Job specializations:
  • Software Development
    Software Engineer, C++ Developer
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below
Position: Chip Performance Profiling Engineer - Member of Technical Staff

Company :
Infinity
· Team :
Systems / AI Infrastructure
· Location :
San Francisco (on-site)
· Type :
Full-time

The Mission

You can't optimize what you can't measure, and on a fresh accelerator there is usually nothing to measure with - no Nsight, no rocprof, no performance counters anyone has documented how to read. Visibility today is a per-vendor artifact, hand-built by the people who shipped the silicon, so every chip without a mature profiler leaves engineers optimizing in the dark until someone ports one over by hand.

And even where a profiler exists, most of them hand you data instead of an answer: a thousand numbers that never say which one is the bottleneck.

We're building the agent that generates that visibility. Give it a supported chip and it produces a profiler for that chip - one that attributes runtime to every operation and sub-operation on an inference pass, fine‑grained enough to show which step inside an attention kernel is the problem rather than just that attention is slow. It reimplements the surface area engineers already expect from nvprof and Nsight, so moving to a new accelerator doesn't cost you the tools you profile with, and it runs on the chip itself with no simulator in the loop.

The same machinery profiles anything the chip does, not only inference.

Measurement sits upstream of everything else this stack does. The optimizer can't move a number it can't see, and an engineer can't fix a bottleneck no tool will name. Making that visibility something we generate rather than something each vendor hand-builds means every accelerator we support arrives already observable - and the profiler points at the specific thing standing between the current code and more performance instead of leaving you to find it in the noise.

What

you’ll work on

You’ll build the instrument the rest of the stack reads from. Depending on your strengths, you’ll own one or more parts of the system:

  • Profiler‑generation agent – the system that takes a supported chip and builds a working profiler for it, so visibility on a new accelerator is something you generate rather than something someone ports by hand every time.
  • Per‑operation attribution – breaking an inference pass down until every operation and sub‑operation carries its own runtime, fine enough to see which step inside a kernel is the one costing you rather than just which kernel is slow.
  • The nvprof and Nsight surface area, per chip – reimplementing the features engineers already lean on to profile, on accelerators that never shipped a profiler of their own, so the tooling doesn’t reset every time the hardware does.
  • Counter and telemetry discovery – finding and validating the signals on parts where the performance‑monitoring unit is undocumented or has to be inferred from behavior, since everything downstream depends on trusting what those counters report.
  • Faithful instrumentation – keeping the act of measuring from changing the timing it measures, because a profiler that perturbs the numbers it reports is worse than no profiler at all.
  • Bottleneck surfacing – turning a wall of measurements into the specific thing standing between the current code and more performance, so both the optimizer and the engineer know where to push.
What we’re looking for

We care more about depth and range than a specific checklist, but strong candidates will have most of:

  • Real performance‑analysis experience – you’ve profiled hard problems and know the difference between a number and a number you can trust.
  • Low‑level systems background – hardware counters, tracing, sampling, and instrumentation.
  • Comfort building measurement tools when the documentation for what you’re measuring is thin or absent.
  • Statistical care – you worry about the observer effect, noise, and sample size before you report a result.
  • Fluency in Python and at least one systems language (Rust, C, or C++).
Nice to have
  • Built a profiler, tracer, or telemetry pipeline that other people relied on.
  • Know the internals of perf, VTune, or Nsight rather than just their front ends.
  • Worked directly with hardware performance counters and their sharp edges.
  • Hands‑on experience building with coding agents.
Who We Are

Infinity is an early‑stage AI infrastructure research company building the software layer that makes non‑NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low‑level code that determines how efficiently a chip runs AI models. We’ve signed or are negotiating design partnerships with d‑Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others.

Founded by Jeremy Nixon (former Google Brain; co‑founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We’re headquartered in San Francisco.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary