×
Register Here to Apply for Jobs or Post Jobs. X

Vice President - AI Safety Platform Engineering

Job in New York, New York County, New York, 10261, USA
Listing for: Goldman Sachs Group, Inc.
Full Time position
Listed on 2026-09-05
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 130000 - 250000 USD Yearly USD 130000.00 250000.00 YEAR
Job Description & How to Apply Below
Location: New York

Vice President - AI Safety Platform Engineering

New York, NY, United States

Job Description

Role Overview

We are seeking a
Vice President – AI Safety Platforms
to build and lead our enterprise AI safety engineering initiatives. As generative AI in financial services evolves from simple prompt-response workflows to autonomous agentic systems that execute multi-step plans, call APIs, and interact directly with internal systems,establishing robust safety mechanisms and standardized evaluation protocols is essential.

In this role, you will recruit and lead dedicated engineering pods focused on developinga unified company-wide agentic evaluation framework
, real-time LLM guardrail services, and automated governance controls. As the senior technical authority for AI safety, you will collaborate closely with core AI platform teams, risk control functions, and business units to drive necessary enhancements to the core AI platform (such as telemetry hooks, API capabilities, execution sandboxes, and data logging infrastructure) to ensure all enterprise AI deployments operate safely, verifiably, and in compliance with institutional standards.

Key Responsibilities

1. Unified Agentic Evaluation Framework

  • Company-Wide Architecture:Design, build, and deploya single, company-wide agentic evaluation framework
    that standardizes how teams across all business lines benchmark, test, and measure AI agent performance prior to production deployment.
  • Trajectory & Multi-Step Reasoning Assessment:Implement evaluation methodologies that score autonomous planning quality, tool-calling precision, multi-turn state retention, trajectory efficiency, and error-recovery behaviors.
  • Continuous Monitoring & Production Drift:Integrate automated evaluation pipelines into runtime environments to continuously audit agent execution traces, detecting reasoning drift, tool failure modes, and unexpected trajectory shifts in production.
  • Domain-Specific Benchmarking:Establishstandardized test suites and synthetic evaluation benchmarks tailored to complex financial workflows, such as automated research, risk assessment, and operational task execution.

2. LLM Guardrails Infrastructure & Real-Time Controls

  • Low-Latency Guardrail Engine:Architect and scale enterprise guardrail microservices that inspect prompt inputs, retrieved context, and model outputs in real time to prevent data leakage, policy violations, and unvalidated execution.
  • Tool-Use & Action Control:Implement runtime policy gateways that inspect and authorize tool calls before execution, ensuring agentsoperatewithin authorized data boundaries and action scopes.
  • Human-in-the-Loop (HITL) Triggers:Build configurable escalation workflows and approval gates that automatically pause execution for high-risk operations (e.g., money movement, client record modifications, or external communications) until human authorization is granted.

3. Core AI Platform Enhancements & Governance Integration

  • Drive Platform Enhancements:Partner directly with the core AI Platform team to drive the implementation of safety APIs, telemetry hooks, developer SDKs, andMLOps/LLMOpspipeline integrations.
  • Auditability & Execution Telemetry:Define and enforce technical standards for immutable audit logging, execution tracing (e.g.,Open Telemetrystandards), and principal identity propagation across all agentic workflows.
  • Regulatory & Model Risk Alignment:Translate model risk management standards (e.g., SR 11-7 / SR 26-2 guidance, FINRA supervision requirements) into automated engineering safeguards and policy checks.

4. Engineering Leadership & Strategic Oversight

  • Team Building & Mentorship:Hire, develop, and mentor high-performing engineering teams specializing in applied machine learning, AI safety, and enterprise platform engineering.
  • Strategic

    Roadmap:

    Own the technical roadmap for enterprise AI safety infrastructure, setting clear milestones for evaluation framework adoption, runtime latency optimization, and governance automation.
  • Stakeholder

    Collaboration:

    Articulate technical risk profiles,evaluationmetrics, and safety architecture to risk committees, model validation teams, and executive leadership.

Key Qualifications

Basic…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary