Research Intern: Meta-Cognition and Internal Mechanisms Multi-Modal Foundation Model- Agent
Listed on 2026-09-13
-
Research/Development
Research Scientist
Research Intern:
Meta-Cognition and Internal Mechanisms for Multi-Modal Foundation Model-based Agent - Honda Research Institute USA Research Intern:
Meta-Cognition and Internal Mechanisms for Multi-Modal Foundation Model-based Agent
Job Number: P25IN
T-60
Honda Research Institute USA (HRI-US) is seeking a self-motivated research intern to join our Cooperative Cognition Team in advancing the next generation of adaptive and meta-cognitive AI agents built on multimodal foundation models. As an intern, you will investigate the internal mechanisms of multimodal foundation models and explore how they can be leveraged to enable agents to monitor, evaluate, and adapt their own reasoning and behavior.
The research will focus on understanding and utilizing internal model representations, such as hidden states, neural activations, attention patterns, and latent structures, to support capabilities including self-monitoring, uncertainty estimation, error detection, self-correction, and adaptive decision-making. The work may also explore these ideas in embodied and robot-learning settings, including learning from video demonstrations of human activities and procedural tasks. In this context, meta-cognitive mechanisms can help an agent assess the quality of its understanding, recognize uncertainty or failure, and adapt its behavior at inference time.
Through this research, you will contribute to developing robust and self-adapting multimodal agents that can better understand their own capabilities and limitations and operate effectively in complex, evolving environments.
San Jose, CA
Key Responsibilities
Over the course of the internship, the intern will contribute to the development of meta-cognitive methods and internal mechanisms that improve the adaptability, efficiency, and reliability of multimodal foundation model-based agents. The research will explore how internal model representations and inference-time mechanisms can be understood and leveraged to enable more capable, self-aware, and adaptive agentic AI systems. Potential research directions include (but are not limited to):
- Developing methods to analyze, probe, interpret, and manipulate internal representations of multimodal foundation models, including hidden states, neural activations, latent features, and computational circuits, to enable meta-cognitive capabilities such as self-assessment, uncertainty estimation, error detection, and adaptive intervention.
- Investigating inference-time and test-time scaling strategies for multimodal foundation models, including vision-language and vision-language-action models, to improve reasoning, planning, adaptation, and self-correction.
- Developing meta-cognitive mechanisms that leverage internal model signals to dynamically assess model behavior, identify limitations or failures, and adapt reasoning or decision-making strategies at inference time.
- Designing and curating benchmarks and evaluation methodologies for assessing meta-cognitive capabilities, internal-state awareness, robustness, and adaptation in multimodal agentic systems across complex and long-horizon tasks.
Minimum Qualifications
- Currently enrolled as a Ph.D. student in Computer Science, Machine Learning, or a related field at a reputed university (exceptional M.S. candidates with a minimum of 1 year of research experience may also be considered.
- Strong familiarity with agentic AI systems, modern foundation models, and their internal representations (e.g., hidden states, neural activations, embeddings, and latent structures).
- Experience in open-source deep learning frameworks
Bonus Qualifications
- Experience with representation learning, mechanistic interpretability, or analysis of neural activations in foundation models.
- Familiarity with techniques for probing, interpreting, or steering internal model behavior (e.g., activation-level analysis, feature attribution, or circuit-level analysis).
- Experience with robot learning, learning from demonstrations, imitation learning, reinforcement learning, multimodal foundation models, or vision-language-action models.
- Familiarity with policy representations, skill learning, action representations, hierarchical decision-making, or grounded multimodal reasoning.
Years of Work Experience Required 0
Desired Start Date 1/11/2027
Internship Duration 3 Months
Position Keywords Agentic AI, Meta-Cognitive AI
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).