Code Quality Engineer LLM Evaluation
Listed on 2026-10-02
-
Software Development
Software Engineer, Software Testing, AI Engineer (Applied/Software)
The Work
You will review AI-generated code used to train and evaluate large language models. Your work will cover different programming languages and software situations, including bug fixes, new features, refactoring, API integrations, configuration changes, and database operations.
You will assess whether code works as intended and identify defects, logic errors, missing pieces, edge-case failures, performance problems, security risks, and architecture weaknesses. You will compare possible implementations, improve code into reliable reference solutions, explain your recommendations, and help create technical rubrics and coding benchmarks.
- Review AI-generated code and software changes for correctness, security, reliability, scalability, readability, and maintainability.
- Debug complex codebases and identify the root causes of technical problems.
- Compare implementations and select or create the solution that best meets the technical requirements.
- Write clear technical feedback and contribute to evaluation criteria, coding benchmarks, and improved LLM evaluation methods.
The role details do not list a pay rate. This is a remote, part-time contract role for an individual contributor supporting AI training and engineering evaluation work.
- Pay:
Not provided in the role details. - Time: 20+ hours per week.
- Location:
Worldwide and remote. - Language:
Written English proficiency is required. - Experience:
The listing is marked entry level in the source fields, while the role description requires at least seven years of professional software engineering experience. - Programming:
Strong proficiency in at least one of Python, JavaScript, Type Script, Java, C++, Go, C#, Ruby, PHP, or Rust. - Technical knowledge:
Production software development, debugging, code review, clean code, modular architecture, abstraction, error handling, data structures, algorithms, APIs, databases, and application architecture. - Practices:
Experience with collaborative code reviews, Git, and modern software engineering methods. - Helpful background:
Experience evaluating AI-generated code, creating technical rubrics, contributing to software engineering benchmarks, or working across multiple languages and architecture patterns.
AI training is the human work behind systems that generate and understand code, text, images, and other data. People review examples, rate model outputs, and provide clear corrections so AI systems become more useful and reliable;
Open Train helps people find and build careers in this field.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).