×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff, Data Infrastructure

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Inception
Full Time position
Listed on 2026-08-18
Job specializations:
  • Software Development
    Data Engineering, Machine Learning/ ML Engineer, Data Scientist
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below

Role

We seek experienced engineers to architect and scale the core infrastructure behind distributed training pipelines and petabyte-scale data catalogs. You ll work directly with researchers to accelerate experiments, develop new datasets, improve infrastructure efficiency, and enable key insights across our data assets.

Key Responsibilities
  • Design, build, and operate scalable, fault-tolerant infrastructure for LLM research: distributed compute, data orchestration, and storage across modalities.
  • Develop high-throughput systems for data ingestion, processing, and transformation — including training data catalogs, deduplication, quality checks, and search.
  • Build systems for web crawling, data ingestion, and real-time data processing to support model training operations.
  • Develop tools and frameworks for efficient data storage, retrieval, and versioning across distributed systems.
  • Ensure data collection adheres to privacy regulations.
Qualifications
  • BS/MS/PhD in Computer Science, Machine Learning, or a related field (or equivalent experience).
  • 3+ years of experience building data processing pipelines at scale, particularly with AI/ML applications.
  • Strong proficiency in Python and experience with data processing frameworks (Apache Spark, Beam, Airflow).
  • Familiarity with synthetic data generation techniques and data augmentation strategies.
  • Familiarity with web scraping, crawling technologies, and Common Crawl datasets.
  • Solid understanding of machine learning fundamentals and experience with ML frameworks (PyTorch, Tensor Flow).
  • Experience with SQL and No

    SQL databases for managing structured and unstructured data.
Preferred Skills
  • Experience with large language models and understanding of tokenization, embeddings, and model architectures.
  • Experience managing human annotation workflows and quality control processes.
  • Experience with vector databases and embedding-based retrieval systems.
  • Knowledge of data privacy regulations and ethical AI practices.
  • Experience with distributed computing and large-scale data storage systems (HDFS, S3, Big Query).
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary