×
Register Here to Apply for Jobs or Post Jobs. X

Manager; AI Evaluation Engineering

Job in Chicago, Cook County, Illinois, 60290, USA
Listing for: Caterpillar Brazil
Full Time position
Listed on 2026-09-04
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), DevOps, Cloud Engineer - Software, Software Project Mgr/ Lead
Salary/Wage Range or Industry Benchmark: 147760 - 240110 USD Yearly USD 147760.00 240110.00 YEAR
Job Description & How to Apply Below

Career Area:
Technology, Digital and Data

Job Description:

Your Work Shapes the World at Caterpillar Inc. When you join Caterpillar, you're joining a global team who cares not just about the work we do - but also about each other. We are the makers, problem solvers, and future world builders who are creating stronger, more sustainable communities. We don't just talk about progress and innovation here - we make it happen, with our customers, where we work and live.

Together, we are building a better world, so we can all enjoy living in it.

Job Summary

Join the AI Engineering team of Cat Digital and take charge of leading a team dedicated to evaluating and validating our advanced generative AI solutions including intelligent agents, digital assistants, and other innovative capabilities. You will shape the development of cutting-edge products that redefine how customers and dealers interact with technology through AI-driven experiences.

What You Will Do
  • Providing strong technical support and clear direction to ensure the team is aligned with company goals and capable of evaluating advanced AI projects.
  • Overseeing the performance of both individual team members and the team as a whole, fostering a culture of learning by identifying and addressing training and development needs.
  • Taking ownership of the quality of AI engineering products, making sure all solutions are robust, reliable, and meet high standards.
  • Establishing and supervising the implementation of engineering best practices to maintain consistency and excellence in development processes.
What You Will Have
  • Products and Services:
    Knowledge of major products and services and product and service groups; ability to apply knowledge of product and service appropriately to diverse situations.
  • Software Development:
    Knowledge of software development tools and activities; ability to produce software products or systems in line with product requirements.
  • Software Development Life Cycle:
    Knowledge of software development life cycle; ability to use a structured methodology for delivering and managing new or enhanced software products to the marketplace.
  • Software Quality Assurance and Testing:
    Knowledge of software quality assurance and testing; ability to apply appropriate processes, tools, and techniques for assuring a high level of quality in computer software products and systems.
Considerations For Top Candidates
  • Proven experience leading software quality, test automation, validation, and AI evaluation teams, with the ability to scale processes, tools, metrics, and engineering practices across multiple products and enterprise initiatives.
  • Demonstrated success delivering enterprise-scale software and GenAI solutions across hybrid cloud and embedded/edge environments, leveraging modern software engineering practices including CI/CD, automated testing, incident management, feature flags, and progressive deployment strategies (e.g., blue/green and canary deployments).
  • Deep expertise in GenAI architecture and system design, including LLMs, SLMs, multimodal, speech, and real-time models; prompt engineering; agentic systems; tool use and orchestration; RAG architectures; vector databases, embeddings, and chunking strategies; fine-tuning techniques such as LoRA; and emerging standards including MCP and A2A.
  • Hands-on experience with AI development and evaluation ecosystems, including Azure, AWS, GCP, Azure AI Foundry, Sage Maker, Bedrock, Snowflake Cortex, Lang Chain, Lang Graph, Langfuse, Arize, Lang Smith, Humanloop, Ragas, Deep Eval, Phoenix, or similar technologies.
  • Deep understanding of AI evaluation methodologies and the challenges of non-deterministic systems, including metric design, RAG quality assessment, output reliability, safety evaluation, A/B testing, human-in-the-loop validation, and balancing deterministic and probabilistic testing approaches.
  • Expertise in modern AI observability and monitoring practices, including tracing, telemetry, evaluation data collection, root-cause analysis, and the use of AI technologies to enhance testing through automated test generation, synthetic data creation, scenario expansion, and LLM-assisted evaluation.
Summary Pay…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary