More jobs:
Principal AI Architect/Engineer
Job in
Plano, Collin County, Texas, 75086, USA
Listed on 2026-08-14
Listing for:
Pepsico
Full Time
position Listed on 2026-08-14
Job specializations:
-
Software Development
DevOps
Job Description & How to Apply Below
The AI Platform/Observability Architect is an execution-focused engineer who designs, builds, and operates observability capabilities within a defined domain of the enterprise AI observability platform. Working under the strategic direction of the Senior AI Observability Architect (L11), this role translates architecture blueprints into production-grade instrumentation, telemetry pipelines, dashboards, quality gates, and safety signals across agentic AI systems.
The junior architect is a hands-on engineer who codes, integrates, tests, and iterates - owning feature-level delivery within one or more specialization tracks while developing a growing understanding of the full observability platform. They are a technical practitioner first, with an emerging architect mindset.
Responsibilities
Observability Platform Engineering & OTEL Integration (25%):
- Implement Open Telemetry (OTEL) instrumentation within assigned agent frameworks or platforms - including custom exporters, span enrichers, semantic conventions, and context propagation hooks.
- Build and maintain telemetry pipeline components (collectors, processors, exporters) that route metrics, logs, traces, and semantic signals to central observability backends.
- Integrate OTEL with enterprise agentic platforms as assigned - which may include Salesforce Agent Force, Service Now, Microsoft Agent 365, or internal frameworks - following architecture blueprints set by the L11.
- Develop and maintain observability dashboards, alerting rules, and SLO/SLA definitions for the assigned sub-domain, ensuring signal quality and low false-positive rates.
- Participate in on-call rotations and incident response for the observability platform - contributing to RCA documentation and runbook improvement.
- Write unit, integration, and end-to-end tests for all telemetry components; maintain >80% test coverage across owned services.
- Instrument safety-critical signal capture within assigned pipelines - including guardrail trigger rates, policy violation events, prompt injection detections, and hallucination flags.
- Support red team exercises by building observability hooks that capture adversarial test results, attack surface telemetry, and behavioral deviation signals in real time.
- Implement secure trace handling for sensitive AI decision events - applying data masking, PII redaction, and audit-log retention policies as defined by the security architecture.
- Assist in maintaining the Security Observability Playbook - documenting findings, updating escalation paths, and contributing to incident classification procedures.
- Monitor agent-to-agent protocol traffic (A2A, UCP, AP2) for anomalous communication patterns and flag deviations for review by the L11 architect and security team.
- Implement RAI signal collectors within assigned agent workflows - capturing fairness indicators, bias detection outputs, explainability scores, and content safety classifications.
- Maintain RAI telemetry pipelines and ensure data quality, completeness, and timeliness of governance signals feeding into compliance dashboards.
- Contribute to audit-readiness work by ensuring all AI decision traces within the assigned domain include required governance metadata and are retained per policy.
- Support gap analyses by comparing current RAI signal coverage against governance framework requirements and flagging coverage gaps to the L11.
- Build and maintain quality gate components within CI/CD pipelines - using observability data to detect performance regressions, behavioral drift, and SLA breaches before they reach production.
- Instrument and monitor Skill Evaluations (evals) across the Memory, Skills, and MCP harness stack - collecting eval results, tracking pass/fail trends, and alerting on regression thresholds.
- Implement continuous quality monitoring for post-go-live agentic solutions - tracking agent success rate, tool-call fidelity, latency distributions, and user outcome proxies.
- Conduct structured testing of new agent capabilities…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×