RL Environment Data Engineer/Researcher Intern
Job in
Greater London, London, Greater London, W1B, England, UK
Listed on 2026-10-04
Listing for:
Eigent
Apprenticeship/Internship
position Listed on 2026-10-04
Job specializations:
-
Research/Development
Research Assistant/Associate, Information & Knowledge Management
Job Description & How to Apply Below
Location: Greater London
RL Environment Data Engineer / Researcher Intern
We are looking for an RL Environment Data Engineer / Researcher Intern to support our team in building and refining reinforcement learning training environments across different domains. Working alongside our researchers and engineers, you will assist with data collection, task definition, reward design, evaluation, and checking how well environment data works in post-training. This internship suits students and early-career researchers who want hands-on experience with RL environments and LLM post-training.
- Assist in designing and improving RL training environments across various task domains.
- Help collect, clean, structure, and evaluate data used for RL environment construction and model post-training.
- Support the team in defining task objectives, reward functions, and evaluation standards.
- Help identify loopholes in reward design and test approaches that prevent reward hacking.
- Run experiments in validation environments and report on the effectiveness of post-training data and environment design.
- Work with research, engineering, and data team members to improve environment coverage, task difficulty, and evaluation reliability.
- Keep up with research on RL environments, data evaluation, AI agents, and post-training methods, and share relevant findings with the team.
- Currently pursuing or recently completed a degree in Computer Science, AI, or a related field.
- Good coding skills in Python, with the ability to build scripts, data pipelines, and simple evaluation tools.
- Comfortable using AI coding tools for code generation, debugging, and rapid experimentation.
- Foundational understanding of reinforcement learning, post-training, reward design, or data evaluation, through coursework, research, or projects.
- Interest in turning real-world tasks into trainable and measurable RL environments.
- Experience with data scraping, data cleaning, annotation, or data quality assessment is a plus.
- Exposure to LLM agents, RLHF/RLAIF, coding agents, automated evaluation, or benchmark construction is a plus.
- Curious and eager to learn, with the ability to iterate quickly based on feedback and experiment results.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×