More jobs:
Principal Core Infrastructure Engineer - AI-Powered Digital Twin Platform
Job in
Nashville, Davidson County, Tennessee, 37201, USA
Listed on 2026-09-01
Listing for:
Oracle
Full Time
position Listed on 2026-09-01
Job specializations:
-
Software Development
Cloud Engineer - Software, Software Architect, Backend Developer, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Description
Oracle is seeking a Principal Software Developer to help design and build an AI-powered Digital Twin platform. The platform will combine real-time operational data, digital models, simulation, graph technologies, and AI agents to help customers understand complex systems, predict outcomes, investigate issues, and automate actions safely.
This is a senior individual contributor role for an engineer who can lead complex technical initiatives, influence architecture across teams, and deliver production-grade capabilities spanning distributed systems, data platforms, generative AI, and agentic workflows.
Responsibilities
Responsibilities
- Design and build scalable services for digital models, telemetry, state management, relationships, simulation, and real-time analytics.
- Develop AI agents that reason over digital twin data, enterprise knowledge, telemetry, and operational systems.
- Build capabilities for RAG, graph-based retrieval, agent orchestration, tool use, memory, evaluation, and human approval.
- Define technical designs and influence architecture across multiple components and engineering teams.
- Develop secure, reliable, and observable distributed systems for enterprise cloud environments.
- Ensure AI-generated recommendations and actions are grounded, explainable, auditable, and policy-compliant.
- Partner with product managers, applied scientists, architects, and engineers to convert customer requirements into scalable platform capabilities.
- Lead complex feature development from design through implementation, testing, deployment, and production support.
- Improve engineering quality through design reviews, code reviews, automated testing, operational readiness, and technical mentoring.
- Investigate and resolve difficult performance, reliability, data consistency, and production issues.
- 8+ years of software engineering experience building large-scale production systems.
- Strong programming skills in Java, Python, Go, C++, or a comparable language.
- Deep understanding of distributed systems, data structures, algorithms, APIs, databases, and cloud architecture.
- Experience designing and operating highly available, multi-tenant cloud services.
- Experience with generative AI, large language models, RAG, AI agents, or agentic orchestration.
- Familiarity with graph databases, vector search, event streaming, time-series data, or simulation systems.
- Experience delivering complex projects across multiple services or teams.
- Strong debugging, problem-solving, and system-design skills.
- Ability to communicate technical decisions clearly and influence without direct authority.
- Experience with digital twins, IoT, industrial systems, cloud operations, or simulation platforms.
- Experience with Lang Graph, Lang Chain, Llama Index, Semantic Kernel, OCI Generative AI Agents, or similar technologies.
- Knowledge of LLMOps, agent evaluation, prompt and policy versioning, model monitoring, and AI governance.
- Experience building knowledge graphs, hybrid retrieval systems, or enterprise RAG solutions.
- Experience applying AI agents to root-cause analysis, predictive maintenance, optimization, incident response, or automated remediation.
- Experience with OCI, Oracle Database, Kubernetes, graph databases, streaming platforms, or vector databases.
- Experience mentoring engineers and providing technical leadership for major platform initiatives.
System Design & Architecture - System Scalability:
-Lead the development and implementation, and begin to architect, components of scalable distributed systems that support horizontal and vertical scaling to meet system demands, including leveraging distributed state management tools.
-Optimize code and/or systems for large-scale data processing and high-throughput requirements to support hyper-scale systems.
-Define scalability requirements for owned components and ensure design and implementation requirements are met.
-Design systems to scale with elasticity (e.g., effectively scaling both up and down).
-Leverage data plane platforms to effectively handle large-scale data retrieval, storage, and processing.
-Design performance and load testing.
System Design & Architecture - System Reliability Design:
-Build and design fault-tolerant components and systems capable of withstanding in-service updates by implementing redundancy, replication, and automatic failover mechanisms.
-Design systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
-Implement and optimize approaches to handle network unreliability, including load-shedding, throttling, and rate-limiting.
-Design components and systems that are durable and adhere to service level objectives (SLOs), setting expectations for availability and durability of other computing services within the department.
System Design & Architecture - System…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×