Staff Data Engineer
Job Description
Own the data flywheel behind physical AI at LiveView Technologies. This onsite Staff Data Engineer role in Seattle focuses on turning raw edge telemetry and video into governed, versioned datasets that enable model training and evaluation. You will partner with AI/ML research and MLOps, setting the standards for how data is stored, labeled, validated, and served so teams work from one consistent source of truth.
What you’ll deliver
- Own the end-to-end loop converting raw edge telemetry and video into labeled training data, frozen evaluation/benchmark sets, and the dataset outputs that feed the next model iteration.
- Build and own pipelines that register raw sources, standardize them into a single well-defined schema, and join and aggregate curated datasets so every team trains, validates, and benchmarks using one consistent store through one reader.
- Append labels and semantic annotations to datasets without rewriting source data, then version, quality-check, and serve the results. Partner with annotation and data-operations teams on label production and verification while you own dataset, storage, and serving.
- Own frozen, versioned validation and benchmark datasets that keep model comparisons stable over time, including the review and scrubbing discipline required before sharing externally.
- Set up schema and content versioning so producers can evolve datasets without breaking consumers via opt-in versions and append-without-rewrite for new fields, supported by reader/writer indirection for controlled rollouts.
- Develop the read/write libraries and integrations researchers depend on, including PyTorch/Lightning dataloaders, a simple record-level CRUDL API, and Spark/analytics access with self-service patterns for AI teams.
- Implement governance machine-enforced throughout the flywheel: classification rules, access control, scrubbing and anonymization in load jobs, and lineage and provenance for each dataset version, annotation campaign, and training input.
- Define data-engineering standards for schema conventions, dataset contracts, and quality gates, and mentor ICs toward those goals while growing the function as the team forms.
What you bring
- 8+ years building and operating large-scale data pipelines and production data-lake or lakehouse systems for ingestion, ETL/ELT, partitioning, storage-format decisions, and consumer-ready reader/writer libraries.
- Experience building pipelines for model training and evaluation, labeled data, and evaluation/benchmark sets, with an understanding of how data quality and versioning impact model results.
- Strong background with medallion-style layered architectures and modern table/lake formats such as Iceberg, Delta, Parquet (or comparable), including schema evolution and dataset versioning.
- Hands-on work with storage and access patterns for large multimodal video, image, and sensor/telemetry data, including queryability at scale (denesting, repartitioning, and binary-inline vs. reference storage).
- Practical experience with the data side of PyTorch/Lightning and Spark, plus strong Python knowledge.
- Practical experience enforcing data governance in pipelines, including classification, access control, lineage and provenance, and retention for privacy-sensitive data.
- A track record of setting data-engineering direction and leveling up engineers (technical leadership; formal management not required).
- Bachelor’s or Master’s in Computer Science, Engineering, or a related field, or equivalent practical experience.
Tools you’ll work with
- PyTorch, Lightning, Spark, Python
- Iceberg, Delta, Parquet
- Kafka, EMR, MCP
- PyTorch/Lightning dataloaders, Encord, Labelbox
- Lambda, CRUDL API
Benefits
- Comprehensive health, dental and vision coverage
- Retirement benefits (401k match up to 4%)
- Flexible PTO
Preferred qualifications
- Streaming or near-real-time ingestion from edge/IoT sources into a data lake (for example Kafka, Lambda, EMR, or similar).
- Append-without-rewrite and hash-indexed dataset techniques on open table formats, plus dataset/feature-versioning systems.
- Generative-AI data work: fine-tuning and evaluation dataset curation for LLMs/VLMs.
- Exposing datasets to AI agents through MCP-style query interfaces, with semantic schema and plain-language documentation for retrieval.
- Computer-vision/video annotation tooling and workflows (for example Encord, Labelbox, or similar).
Compensation
The beginning annual salary range for this role is $171,900 - $221,000 USD, determined by location, job-related experience, and education/training. Total earning potential is amplified by a bonus structure tied to meeting goals, and you will become an owner from day one through the employee equity program.
Location: Seattle, WA (onsite)