More jobs:
Engineering Expert
Job in
San Diego, San Diego County, California, 92189, USA
Listed on 2026-10-02
Listing for:
Turing Global India
Full Time, Part Time
position Listed on 2026-10-02
Job specializations:
-
Engineering
AI Evaluation, Systems Engineer
Job Description & How to Apply Below
Role Overview
We are seeking experienced AI Evaluation Engineers (Engineering Simulation & Design) to author and validate "model-breaking," simulation-based engineering design problems to train and evaluate state-of-the-art AI agents. Operating across major engineering disciplines including Electrical, Mechanical, Control Systems, Aerospace, Systems, and Robotics—you will create complex, multi-constraint tasks where AI agents must interpret requirements, navigate trade-offs, configure open-source simulation tools, diagnose failures, and iterate toward valid solutions.
You will analyze agent execution logs, expose systemic reasoning gaps, and build automated, objective graders to elevate frontier model performance.
- Education & Expertise: Master’s degree or PhD in Electrical, Mechanical, Aerospace, with 10+ years of hands-on engineering design experience.
- Simulation Tooling: Proficiency with at least one domain-relevant open-source simulation package (e.g., ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, Cad Query, build
123d, Open Modelica, Cantera, Gmsh) combined with strong Python scripting skills. - AI Evaluation & Failure Diagnostics: Hands-on experience with modern LLMs/coding agents and evaluation concepts (pass@k, failure-mode analysis, nondeterministic behavior), with the ability to audit trajectory logs and isolate core reasoning/tool-use failures.
- Domain Rigor & Precision: Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.
- Availability & Commitment: Talent must have weekend on-call availability (part-time engagement is acceptable).
- Technical Infrastructure: Personal desktop/laptop equipped with a stable, high-speed internet connection in a remote setup.
- Model-Breaking Problem Design: Author original, self-contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective autograders.
- Environment & Simulation Integration: Build, run, and validate problem environments using open-source simulation tools and custom Python test benches.
- Trajectory Analysis & Failure Mode Taxonomy: Evaluate coding agent outputs and execution logs across repeated trials to identify systemic failure modes (e.g., misinterpreting simulator feedback, premature design convergence, physically impossible geometries).
- Difficulty Calibration & Benchmark Refinement: Iteratively refine problem difficulty based on empirical model performance data without introducing ambiguity or missing information.
- Cross-Functional Collaboration: Partner with AI researchers, pod leads, and domain experts to integrate high-rigor benchmarks into the model evaluation pipeline.
- Electrical Engineering
- Mechanical Engineering
- Aerospace Engineering
- Master or PhD from relevant engineering domain is a must for this position.
- Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role is required.
- Commitments
Required:
40 hours per week with 4 hours of overlap with PST. - Engagement type:
Contractor - Engagement Length: upto 24 weeks
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×