Machine Learning Engineer
Job Description
Samsung Austin Semiconductor is hiring a Machine Learning Engineer to build and maintain end-to-end model pipelines focused on anomaly detection and root cause analysis. The role emphasizes distributed processing of manufacturing time-series and operational data, with heavy use of PySpark to take datasets from ingestion through training, validation, deployment, and monitoring in production.
This full-time onsite position is located in Taylor, TX and offers a base salary range of $90,000 to $174,500 per year. The team is focused on building reproducible, auditable ML workflows that can be reliably retrained and rolled back when performance shifts.
What you’ll do
- Develop PySpark workflows to ingest, clean, and transform high-volume manufacturing data, producing structured datasets for training and inference.
- Improve Spark performance by tuning partition strategies, managing executor memory, reducing shuffle, and addressing skewed joins to lower runtime and cluster usage.
- Create and maintain end-to-end ML pipelines that automate feature calculation, model training, validation, and deployment while ensuring runs are reproducible and auditable.
- Implement and tune machine learning models for anomaly detection and root cause analysis.
- Own the production model lifecycle including version tracking, secure artifact storage, automated retraining triggers, and rollback procedures when performance degrades.
- Monitor pipeline execution times, data quality, and model metrics such as accuracy, drift, and throughput, and build alerting rules to detect failures or degradation early.
What you bring
- A Bachelor’s degree or higher in Computer Science, Software Engineering, Data Science, or a related quantitative field.
- 3 to 5+ years of professional experience building and maintaining machine learning systems.
- Strong proficiency in PySpark and distributed data processing, including experience optimizing jobs for speed and memory.
- Hands-on experience with Python ML libraries including scikit-learn, TensorFlow, PyTorch, or XGBoost.
- Practical knowledge of MLOps, including pipeline orchestration, model versioning, experiment tracking, and deployment.
- Experience setting up monitoring and alerting for both data pipelines and deployed models.
Technologies
- PySpark, Spark
- Python
- scikit-learn
- TensorFlow
- PyTorch
- XGBoost
Benefits
- Medical, dental, and vision insurance
- Life insurance and 401(k) matching with immediate vesting
- Onsite café(s) and workout facilities
- Paid maternity and paternity leave
- Paid time off (PTO) + 2 personal holidays and 10 regular holidays
- Wellness incentives and MORE
- Eligible full-time employees (salaried or hourly) may receive MBO bonuses based on company, division, and individual performance
Preferred
- Experience setting up model registries, automated retraining triggers, and rollback procedures to keep production models reliable.
- Experience writing automated tests and validation checks for data pipelines and model outputs to catch errors before deployment.
- Familiarity with on-prem or private cloud infrastructure, including cluster management and secure artifact storage.
Additional information
U.S. Export Control Compliance: This role may require access to information subject to U.S. export control laws. Applicants must be authorized to access such information or eligible for government authorization.
Trade Secrets Notice: By submitting an application, you agree not to disclose to Samsung, or encourage Samsung to use, any confidential or proprietary information (including trade secrets) belonging to a current or former employer or other entity.