Test Engineer-AI/LLM
Listed on 2026-08-22
-
Software Development
AI QA / Validation Engineer, AI Engineer (Applied/Software), AI Reliability/ Performance Engineer, Machine Learning/ ML Engineer
OPPO US Research Center is seeking a full‑time meticulous and innovative AI/LLM Test Engineer to join our cutting‑edge AI team. In this critical role, you will evaluate the performance, reliability, and safety of Large Language Models (LLMs) in real‑world product scenarios and test end‑to‑end generative AI solutions. Your work will directly shape how users experience AI‑powered features by ensuring robustness, accuracy, and alignment with product goals.
This is a unique opportunity to pioneer testing methodologies for next‑generation AI systems at the forefront of technology.
Position Requirements Core Testing & Evaluation
- Design and execute performance tests for LLMs across diverse product use cases (e.g., chatbots, content generation etc.)
- Develop automated test frameworks to evaluate LLM outputs for accuracy, bias, safety, and coherence
- Conduct end‑to‑end testing of integrated generative AI solutions, including APIs, data pipelines, and user interfaces
- Collaborate with ML engineers to validate fine‑tuned models and optimize prompts for target scenarios
- Analyze model failures, edge cases, and adversarial inputs to identify risks and improvement areas
- Benchmark LLM performance against industry standards and product‑specific KPIs
- Partner with product, engineering, and research teams to define test requirements and acceptance criteria
- Document defects, performance metrics, and test results to drive data‑driven improvements
- Advocate for AI ethics and safety through rigorous testing of fairness, bias mitigation, and content moderation
- Build scalable tools for synthetic test data generation, prompt variation testing, and automated evaluation workflows
- Stay current with advancements in generative AI testing, including red‑teaming techniques and evaluation frameworks (e.g., HELM, Dynabench)
- Propose novel testing strategies for emerging challenges (e.g., hallucinations, context drift)
- Bachelor’s degree in Computer Science, Data Science, Engineering, or a related technical field, or equivalent practical experience
- 1+ years of experience in software testing, data science, or ML validation, with exposure to AI/ML systems
- Proficiency in Python and testing frameworks (e.g., PyTest, Selenium)
- Hands‑on experience evaluating LLMs in production environments (e.g., GPT, Claude, Llama, Gemini)
- Strong analytical skills for dissecting model behavior, statistical performance, and failure modes
- Familiarity with cloud platforms (GCP, Azure, or AWS) and MLOps tooling (e.g., MLflow, Weights & Biases)
- Experience with version control (Git) and agile development methodologies
- Master’s degree in AI, Machine Learning, or a related field
- Expertise in prompt engineering, LLM fine‑tuning (e.g., LoRA, RLHF), or optimization techniques
- Experience with automated evaluation tools (e.g., Lang Chain, Tru Lens) or LLM‑specific test suites
- Knowledge of data pipelines, SQL/No
SQL databases, and API testing (e.g., Postman) - Background in statistics, quantitative analysis, or data visualization for test insights
- Contributions to AI safety/ethics initiatives or open‑source LLM evaluation projects
- Experience testing mobile‑integrated AI solutions (Android/iOS)
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies, execute evaluation workflows, and assist in model performance validation across diverse generative AI use cases. This contract role is ideal for someone with hands‑on experience in AI/ML evaluation, QA engineering, or data analysis who wants to deepen their exposure to generative AI systems.
ContractorPosition Requirements Testing & Evaluation Support
- Execute pre‑defined performance tests for LLMs across various tasks (e.g., summarization, Q&A, chatbot flows).
- Run scripted evaluations to assess outputs for factuality, coherence, and safety.
- Perform manual and automated test execution on APIs and LLM-integrated user interfaces.
- Assist ML engineers in evaluating…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).