×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

Principal AI Architect​/Engineer

Job in Plano, Collin County, Texas, 75086, USA
Listing for: Pepsico
Full Time position
Listed on 2026-08-14
Job specializations:
  • Software Development
    DevOps
Job Description & How to Apply Below
Overview

The AI Platform/Observability Architect is an execution-focused engineer who designs, builds, and operates observability capabilities within a defined domain of the enterprise AI observability platform. Working under the strategic direction of the Senior AI Observability Architect (L11), this role translates architecture blueprints into production-grade instrumentation, telemetry pipelines, dashboards, quality gates, and safety signals across agentic AI systems.

The junior architect is a hands-on engineer who codes, integrates, tests, and iterates - owning feature-level delivery within one or more specialization tracks while developing a growing understanding of the full observability platform. They are a technical practitioner first, with an emerging architect mindset.

Responsibilities

Observability Platform Engineering & OTEL Integration (25%):
  • Implement Open Telemetry (OTEL) instrumentation within assigned agent frameworks or platforms - including custom exporters, span enrichers, semantic conventions, and context propagation hooks.
  • Build and maintain telemetry pipeline components (collectors, processors, exporters) that route metrics, logs, traces, and semantic signals to central observability backends.
  • Integrate OTEL with enterprise agentic platforms as assigned - which may include Salesforce Agent Force, Service Now, Microsoft Agent 365, or internal frameworks - following architecture blueprints set by the L11.
  • Develop and maintain observability dashboards, alerting rules, and SLO/SLA definitions for the assigned sub-domain, ensuring signal quality and low false-positive rates.
  • Participate in on-call rotations and incident response for the observability platform - contributing to RCA documentation and runbook improvement.
  • Write unit, integration, and end-to-end tests for all telemetry components; maintain >80% test coverage across owned services.
Safety, Security & Red Teaming Support (15%):
  • Instrument safety-critical signal capture within assigned pipelines - including guardrail trigger rates, policy violation events, prompt injection detections, and hallucination flags.
  • Support red team exercises by building observability hooks that capture adversarial test results, attack surface telemetry, and behavioral deviation signals in real time.
  • Implement secure trace handling for sensitive AI decision events - applying data masking, PII redaction, and audit-log retention policies as defined by the security architecture.
  • Assist in maintaining the Security Observability Playbook - documenting findings, updating escalation paths, and contributing to incident classification procedures.
  • Monitor agent-to-agent protocol traffic (A2A, UCP, AP2) for anomalous communication patterns and flag deviations for review by the L11 architect and security team.
Responsible AI (RAI) & Governance Signal Instrumentation (10%):
  • Implement RAI signal collectors within assigned agent workflows - capturing fairness indicators, bias detection outputs, explainability scores, and content safety classifications.
  • Maintain RAI telemetry pipelines and ensure data quality, completeness, and timeliness of governance signals feeding into compliance dashboards.
  • Contribute to audit-readiness work by ensuring all AI decision traces within the assigned domain include required governance metadata and are retained per policy.
  • Support gap analyses by comparing current RAI signal coverage against governance framework requirements and flagging coverage gaps to the L11.
Quality Engineering for Agentic Solutions - Post Go-Live & Continuous QE (15%):
  • Build and maintain quality gate components within CI/CD pipelines - using observability data to detect performance regressions, behavioral drift, and SLA breaches before they reach production.
  • Instrument and monitor Skill Evaluations (evals) across the Memory, Skills, and MCP harness stack - collecting eval results, tracking pass/fail trends, and alerting on regression thresholds.
  • Implement continuous quality monitoring for post-go-live agentic solutions - tracking agent success rate, tool-call fidelity, latency distributions, and user outcome proxies.
  • Conduct structured testing of new agent capabilities…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary