Foundation Model Data Engineer
Job Description
Sciforium is seeking a Foundation Model Data Engineer to shape the data strategy behind its foundation models, from raw acquisition to curated training corpora. In this onsite role in San Francisco, CA, you will own the end-to-end data lifecycle that supports pre-training, post-training, and multimodal alignment at web-scale.
This position focuses on building and operating large-scale pipelines that transform petabytes of unstructured inputs into datasets designed for strong downstream performance, including LLM pre-training mixes, RLHF/DPO preference data, SFT instruction corpora, and vision and video materials.
Responsibilities
- Own end-to-end creation of pre-training datasets for LLMs, including defining the mix of web data, code, books, and technical papers to optimize downstream model performance.
- Design and implement data pipelines for data cleaning, exact and fuzzy deduplication, and high-quality signal extraction from petabytes of raw unstructured data.
- Lead development of post-training datasets, including Supervised Fine-Tuning (SFT) instructions, multi-turn dialogues, and preference modeling data for RLHF/DPO.
- Drive acquisition and processing of vision and video data, addressing multimodal alignment issues, video compression, and temporal data consistency.
- Develop high-throughput data processing scripts in Python, using multiprocessing and multithreading to prevent ingestion and transformation bottlenecks at scale.
- Run statistical deep dives on training corpora to surface biases, knowledge gaps, and quality regressions, with the aim of keeping the model’s training “diet” mathematically balanced.
- (Added Value) Design pipelines that generate high-reasoning synthetic data to fill gaps in natural datasets, using existing models to support labeling and refinement.
Requirements
- 5+ years of industry experience in Data Science or Machine Learning, with a proven track record building and managing datasets for foundation models.
- Expert-level skills in high-performance code, including multiprocessing, multithreading, and efficient memory management for large-scale data tasks.
- Demonstrated experience working with petabyte-scale datasets used to train production-grade LLMs or Large Vision Models.
- Hands-on experience building massive LLM training sets from scratch, including raw web crawls such as Common Crawl and specialized domain data.
- Experience building datasets for RLHF, DPO, and multi-turn instruction following, including managing human-labeling workflows and quality gold-sets.
- Mastery of data-at-scale frameworks such as Spark and Ray, or high-performance data-loading formats such as WebDataset and Parquet.
Technologies
Python, multiprocessing, multithreading, Spark, Ray, WebDataset, Parquet, Common Crawl, RLHF, DPO, SFT
Benefits
- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary and equity
Nice-to-Haves
- Experience building large-scale image or video datasets from scratch (for example, LAION-style pipelines).
- Familiarity with large-scale crawling of multimodal data, including video processing challenges such as codecs and compression.
- Experience designing complex labeling schemas for reasoning, coding, and mathematical benchmarks.
- A Master’s or PhD in a quantitative field with a focus on data-centric AI or information retrieval.
Location: San Francisco, CA (onsite)
Compensation: USD 155,000 - 210,000 per year
Minimum Experience: 5 years