×
Register Here to Apply for Jobs or Post Jobs. X

Principal Data Architect and Manager - Service Special Projects

Job in Cupertino, Santa Clara County, California, 95014, USA
Listing for: Socket.dev
Full Time position
Listed on 2026-08-06
Job specializations:
  • Software Development
    Data Engineering, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 180000 - 260000 USD Yearly USD 180000.00 260000.00 YEAR
Job Description & How to Apply Below

We're building the large-scale data foundation that powers private, personalized experiences across Apple platforms. Our team designs and operates the systems that ingest, unify, and understand information at massive scale — turning petabytes of data from many sources into a single, high-quality, richly structured representation. This foundation is what intelligent search and on-device experiences rely on, and we build it with an uncompromising bar for data quality, freshness, and privacy.

We are looking for a Principal Data Architect and Manager to serve as both the senior technical authority and the people leader for our data platform.

DESCRIPTION

As the Principal Data Architect and Manager on our team, you will serve as both the senior technical authority and the people leader for our data platform. You'll define and own the end-to-end architecture of a real‑time, petabyte‑scale data backbone: from ingestion through a multi‑layered lakehouse to normalized serving layers that power downstream search, ranking, and on‑device experiences. You'll also build, grow, and lead the team of data engineers who bring that architecture to life.

This is a hands‑on principal role with multiple facets: you set the technical vision, personally shape the hardest architectural decisions, drive the roadmap through to production, and manage, mentor, and grow the engineers executing against it. Your leverage comes equally from what you design and from the team you build.

MINIMUM QUALIFICATIONS

MS Degree in Computer Science or related degree and 12+ years of experience inData Architecture, Data Engineering, or Platform Engineering, with at least 5years operating in a Principal, Staff, or Lead Manager capacity. Proven experience leading and managing engineers including hiring, performance management, and technical mentorship of senior ICs and managers. Track record ofshipping petabyte‑scale, low‑latency data platforms in production and operatingthem under real‑world load.

Deep cloud expertise: expert‑level proficiency withcloud object storage (e.g., AWS S3) and its architectural nuances for massivedata lakes and lake‑houses. Experience architecting systems for entityresolution, conflation, or knowledge‑graph construction at scale — ideallyinvolving billions of frequently updated entities. Experience designingpipelines that process multimodal data (structured, text, image) and integrateML model inference including LLMs and embedding models: for enrichment and transformation. Familiarity with LLM/model‑serving infrastructure trade‑offs(inference runtimes, GPU‑backed serving) to inform architectural decisions

Streaming expertise: deep, hands‑on knowledge of Apache Kafka (or comparablebrokers like Kinesis) and complex stream processing (Spark Structured Streaming,Flink, or similar). Data modeling: exceptional ability to design logical and physical data models for large‑scale ingest, retrieval, and analyticalconsumption — including dimensional modeling and lakehouse patterns. Experiencedefining SLAs, quality metrics, and observability standards for large‑scale data platforms, with hands‑on use of monitoring/alerting tooling (e.g.,Prometheus/Grafana, Datadog, or Open Telemetry‑based tracing).

Programming:command of at least one modern data‑pipeline language (Scala, Java, or Python) and strong software engineering fundamentals. Cloud services integration: proven experience wiring together event notifications, queuing, orchestration, andcompute services into resilient production pipelines. Experience with vector search technologies (e.g., Pinecone, Milvus) and storing/serving embeddings(e.g., pgvector, Milvus, FAISS) Excellent written and verbal communication;proven ability to align engineers, partner teams, and senior leadership from multiple lines of business around a shared technical direction, with experience bringing a consumer‑oriented product from inception to production.

PREFERRED

QUALIFICATIONS

Experience with embedding storage and retrieval (e.g., pgvector, Milvus, FAISS) and with graph databases (e.g., Tiger Graph, Neo4j). Experience deploying,serving, and optimizing LLMs or ML models directly in the production, inferenceruntimes/compilers (ONNX Runtime, TensorRT/TensorRT‑LLM), and serving frameworks(Triton, vLLM, Torch Serve or similar). Experience tuning batching, KV‑cache, andGPU utilization for low‑latency, high‑throughput real‑time inference in a data pipeline Experience with data governance tools (e.g., Apache Atlas, AWS Glue Catalog, Data Hub).

Familiarity with Infrastructure as Code (Terraform, Pulumi) and modern CI/CD practice. Experience designing systems that handle petabytes ofunstructured media data. Working knowledge of data privacy regulations and best practice…

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary