×
Register Here to Apply for Jobs or Post Jobs. X

Language Model Inference System Engineer Graduate; Applied Machine Learning

Job in San Jose, Santa Clara County, California, 95199, USA
Listing for: ByteDance
Full Time position
Listed on 2026-09-09
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Software Engineer
Salary/Wage Range or Industry Benchmark: 128000 - 256000 USD Yearly USD 128000.00 256000.00 YEAR
Job Description & How to Apply Below
Position: Large Language Model Inference System Engineer Graduate (Applied Machine Learning) - 2027 Start

Responsibilities

Volcano Ark is an all-in-one large model service platform launched by Volcano Engine. It is a leading platform in China's large model market by product capability and market share. The platform provides end-to-end services including model inference, evaluation, fine-tuning, AI application development, and a plugin ecosystem. Volcano Ark hosts Doubao and leading industry large models, and supports enterprise AI adoption through stable, secure, and trusted solutions as well as professional algorithm and technical services.

Volcano Ark is an all-in-one large model service platform launched by Volcano Engine. It is a leading platform in China's large model market by product capability and market share. The platform provides end-to-end services including model inference, evaluation, fine-tuning, AI application development, and a plugin ecosystem. Volcano Ark hosts Doubao and leading industry large models, and supports enterprise AI adoption through stable, secure, and trusted solutions as well as professional algorithm and technical services.

Data AML is Byte Dance's machine learning platform team. It provides training and inference systems for recommendation, advertising, computer vision, speech, and NLP scenarios across products such as Douyin, Toutiao, and Xigua Video. The team also supports internal business teams with large-scale machine learning compute, explores general and innovative algorithms for business problems, and offers core machine learning and recommendation system capabilities to external enterprise customers through Volcano Engine.

We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.

Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.

Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.

Responsibilities
  • Participate in the engineering development of the Volcano Ark MaaS inference system, optimizing large model inference performance, cost, and stability in ultra-large-scale heterogeneous inference clusters.
  • Reduce large model inference costs through system-level approaches such as disaggregated multi-role inference, distributed KV Cache systems, heterogeneous inference, elastic computing, and multi-tenant co-located inference.
Qualifications

Minimum Qualifications:

  • Individuals who are completing or have recently completed a Bachelor s or Master s degree in Computer Science or a related discipline.
  • Strong command of algorithms, design patterns, and data structures, with solid knowledge of operating systems and computer architecture.
  • Proficient in one or more programming languages such as C++ or Python, with good coding style.
  • Understands GPU hardware architecture, is familiar with high-performance computing software stacks such as CUDA, and has experience in GPU performance analysis.
  • Strong interest in distributed systems and large-scale heterogeneous inference; enjoys studying low-level principles and performance bottlenecks, and actively follows progress in related fields.
Preferred Qualifications
  • Experience in large model inference system optimization, with practical understanding of PD disaggregation, KV Cache systems, and multi-node inference.
  • Experience in distributed network communication optimization, with deep understanding of distributed communication operator implementation and RDMA principles.
  • Familiarity with resource orchestration and scheduling frameworks such as Kubernetes and Ray.
Job Information

For Pay Transparency Compensation Description (Annually)

The base salary range for this position in the selected city is $128000 - $256000 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate s qualifications, skills, competencies and experience, and location.…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary