×
Register Here to Apply for Jobs or Post Jobs. X

Agentic AI, LLM Evaluation, and Systems Research Internship at Confidential Princet

Job in Princeton, Mercer County, New Jersey, 08543, USA
Listing for: Ellenco Estágios e Treinamentos
Apprenticeship/Internship position
Listed on 2026-07-16
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI QA / Validation Engineer, AI Reliability/ Performance Engineer, Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 38572 - 57859 USD Yearly USD 38572.00 57859.00 YEAR
Job Description & How to Apply Below
Position: Agentic AI, LLM Evaluation, and Trustworthy Systems Research Internship at Confidential Princet[...]

Agentic AI, LLM Evaluation, and Trustworthy Systems Research Internship

Siemens invites you to join our Software Systems and Processes team in Princeton, NJ to research and develop scalable intelligent systems using LLMs and semantic technologies. This internship focuses on implementing a Verification and Validation (V&V) framework for multi-agent systems, ensuring the reliability, robustness, safety, and transparency of agentic AI systems in complex, uncertain environments.

Key Responsibilities
  • Research, design, and prototype V&V methods for multi-agent and agentic AI systems, focusing on reliability, safety, repeatability, explainability, and robustness.
  • Develop evaluation harnesses, benchmarks, and test scenarios for LLM-based agents, including tool use, multi-step reasoning, orchestration, failure-mode analysis, and adversarial or edge-case behavior.
  • Implement proof-of-concept prototypes in Python using modern AI and agent frameworks, formal methods, testing technologies, and retrieval-augmented or knowledge-grounded architectures where appropriate.
  • Investigate verification strategies such as model checking, property-based testing, fuzz testing, static or dynamic analysis, runtime monitoring, guardrails, and trace-based observability for complex intelligent systems.
  • Collaborate with researchers and engineers to define milestones, run experiments, analyze results, and translate research insights into scalable industrial software concepts.
  • Document findings, contribute to scientific publications or technical reports, and present results clearly to internal and external technical audiences.
Basic Qualifications
  • Currently enrolled in a PhD program in Computer Science, Artificial Intelligence, Machine Learning, Software Engineering, Formal Methods, or a closely related technical field.
  • 3+ years of research or hands‑on experience in AI, machine learning, generative AI, software engineering, formal methods, autonomous systems, or intelligent agent systems.
  • Strong programming skills in Python and practical experience with modern ML or LLM tooling such as PyTorch, Hugging Face Transformers, Lang Chain, Lang Graph, Auto Gen, Semantic Kernel, CrewAI, or comparable frameworks.
  • Hands‑on experience building, evaluating, or testing LLM‑powered applications, agentic workflows, multi‑agent systems, or AI‑enabled software engineering tools.
  • Strong understanding of software architecture, engineering principles, testing methodologies, experimentation, and empirical evaluation of complex systems.
  • Demonstrated ability to conduct independent research, read and synthesize technical literature, analyze complex problems, prototype solutions, and communicate findings clearly.
  • Proficient in English, both written and verbal.
  • The position requires the applicant to be in the United States of America and hold a valid work permit in the US for the duration of the internship.
Preferred Skills
  • Research experience in formal verification, model checking, theorem proving, runtime verification, AI safety, robust AI, explainable AI (XAI), or trustworthy machine learning.
  • Experience with evaluation of LLMs or agents, including hallucination analysis, benchmark design, tool‑use evaluation, prompt‑injection testing, red teaming, or reliability metrics.
  • Familiarity with RAG architectures, vector databases, knowledge graphs, semantic technologies, ontologies, or graph-based reasoning.
  • Understanding of reinforcement learning, planning, reward modeling, preference optimization, or post‑training approaches for LLMs and autonomous agents.
  • Experience with cloud‑native or distributed systems concepts, microservice architectures, APIs, CI/CD, Git, Docker, Kubernetes, Azure, AWS, or comparable platforms.
  • Experience with testing frameworks for complex software systems, including property-based testing, fuzz testing, simulation-based testing, static analysis, or execution-based evaluation.
  • Track record of research publications, open‑source contributions, academic projects, or demonstrable prototypes related to AI, software engineering, formal methods, or agentic systems.
  • Excellent problem‑solving skills, attention to detail, and ability to quickly learn and apply…
Position Requirements
Less than 1 Year work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary