×
Register Here to Apply for Jobs or Post Jobs. X

Technical AI Evaluation Specialist

Job in Philadelphia, Philadelphia County, Pennsylvania, 19117, USA
Listing for: Innodata India
Full Time position
Listed on 2026-09-02
Job specializations:
  • Software Development
    AI QA / Validation Engineer, Software Testing
Salary/Wage Range or Industry Benchmark: 90000 - 130000 USD Yearly USD 90000.00 130000.00 YEAR
Job Description & How to Apply Below

ROLE OVERVIEW & OBJECTIVE:

We are seeking rigorous Technical AI Evaluation Specialists to benchmark, evaluate, and align frontier Large Language Models (LLMs) specialized in code generation, multi-turn technical reasoning, and software architecture. In this role, you will analyze model-generated

code against strict correctness, complexity, security, and UI fidelity standards across text-to-code, image-to-code, and side-by-side (SxS)

comparison tasks. You will be responsible for uncovering subtle failure modes, edge-case vulnerabilities, and producing evidence-based,

defensible rationales to guide model fine-tuning and Reinforcement Learning from Human Feedback (RLHF).

KEY RESPONSIBILITIES & CORE WORKFLOWS
  • Model-Generated Code Evaluation: Evaluate AI-generated code for syntactical validity, execution accuracy, algorithmic complexity,

and architectural best practices across diverse languages.

  • Image-to-Code & UI Verification: Assess model capability in rendering pixel-perfect, responsive front-end components from

wireframes, mockups, and UI design screenshots.

  • Text-to-Code & Pairwise Analysis: Perform rigorous side-by-side (SxS) evaluations to determine model preference, scoring

completions against granular multi-dimensional rubrics.

  • Failure Mode & Edge Case Reasoning: Stress-test model completions against corner cases, boundary conditions, race conditions,

memory leaks, and input sanitization vulnerabilities.

  • Evidence-Based Rationales: Write authoritative, structured C1-level technical rationales explaining exact point deductions, execution

trace errors, and counterfactual fixes.

  • Guideline Calibration & Feedback: Collaborate with research engineers and prompt authors to refine evaluation rubrics, establish

baseline test harnesses, and identify emerging model degradation patterns.

Mandatory Requirements:
  • Experience:

    2+ years of professional software development

experience OR a strong Computer Science / Software Engineering

degree with demonstrable coding proficiency.

  • Front-End / UI Exposure:
    Hands-on experience translating

design mockups to responsive front-end code (HTML5, modern

CSS, React, or Type Script frameworks).

  • Code Review & RLHF Exposure:
    Proven background in

structured code reviews, automated testing, or prior experience in

AI model evaluation/RLHF pipelines.

  • Language & Communication: C1-equivalent professional English

proficiency with exceptional technical articulation and defensive

writing discipline.

Preferred Qualifications:
  • Multi-Language Fluency: Strong capability in Python,

JavaScript/Type Script, Java, C++, or Go, with an aptitude for

rapidly reading and debugging unfamiliar frameworks.

  • Dev Ops & Tooling: Familiarity with Git, Docker, CI/CD pipelines,

AST parsers, or static code analysis tools.

  • Advanced CS Foundations: Deep understanding of data

structures, algorithms, concurrency, and secure coding practices.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary