Applied Machine Learning Engineer
Artificial Intelligence
Big Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Integration
Data Pipeline
Data Platform
Data Processing
Database
Databases
ETL
Informatica
Information Technology (IT)
Integration
Programming
Programming Language
Programming Languages
SQL
Job Description
Vulcan Elements is hiring an Applied Machine Learning Engineer (onsite) to design and deliver the data infrastructure and pipelines that support operational analytics and AI/ML workloads. The role spans platform selection, data architecture, ETL/ELT development, and data quality practices with required handling for CUI and ITAR data.
Key Responsibilities
- Design and own Vulcan’s data architecture from operational data stores through ETL pipelines to the analytics and AI layer
- Evaluate and select platforms for the data Lakehouse, ETL tooling, and operational databases, considering scalability, compliance needs, operational burden, and cost
- Review, refine, and implement data architecture design documents, ensuring technical soundness and CUI and ITAR data handling requirements are addressed
- Make and document key platform and design decisions with sufficient clarity to support future team members
- Ensure the architecture can scale from pilot plant to full-scale facility without requiring fundamental redesign
- Apply engineering practices across version control, testing, observability, and documentation, and maintain these standards as the data team grows
- Design and build ETL pipelines moving data from operational data stores into the data Lakehouse with contextual enrichment for analytics and AI workloads
- Build dependable ingest paths for structured data, time-series data, files, images, and other outputs from manufacturing and lab systems
- Collaborate with engineering, operations, and IT to understand data flows, dependencies, and integration needs, then translate them into pipeline and architecture decisions
- Identify and remove manual data workflows, replacing them with monitored and reliable pipelines
- Diagnose and resolve data quality issues across the stack, with monitoring designed into pipelines to surface problems early
- Define data models for operational queries, analytical workloads, and future AI and ML applications
- Own data contextualization standards so each data point includes the metadata required to make it meaningful
- Contribute to schema design and payload definitions for operational data stores to improve consistency and legibility
- Support reporting and visibility tools that provide operations and leadership with clear insight into process and quality data
- Write technical documentation covering architecture decisions, data models, pipeline designs, and operational runbooks
Required Qualifications
- 8+ years of experience in data engineering, data infrastructure, or a closely related technical role, with a proven track record of owning and delivering production systems
- Experience designing and building data lakes, Lakehouses, or analytical data stores; ability to weigh platform tradeoffs and defend selection decisions
- Strong experience designing and building ETL/ELT pipelines that enrich and contextualize data
- Deep fluency in data modeling for both operational and analytical workloads, able to design schemas that meet present needs while preserving future options
- Experience with relational databases (PostgreSQL, SQL Server, or similar), with confidence writing and debugging SQL
- Comfort operating in a fast-moving environment with a small team, making decisions with incomplete information and documenting them for future colleagues
- Strong communication skills across technical and non-technical stakeholders, translating operational needs into data architecture decisions
- Must be a U.S. Person due to required access to U.S. export-controlled information or facilities
Technology Stack
- PostgreSQL, SQL Server
- InfluxDB, TimescaleDB
- Delta Lake, Apache Iceberg
- Airflow, Prefect, dbt
- Python, SQL
- AWS, Azure, GCP
- MQTT
- CUI, ITAR, EAR
Desired Skills
- Experience with time-series databases (InfluxDB, TimescaleDB, or similar) used in industrial and IoT contexts
- Familiarity with industrial data concepts such as historian data, process tags, and OT/IT integration, including manufacturing-focused data challenges
- Experience working with or alongside a Unified Namespace or MQTT-based data architecture, and understanding how industrial messaging infrastructure maps to the data layer
- Familiarity with data Lakehouse platforms and open table formats (Delta Lake, Apache Iceberg, or similar)
- Experience with ETL orchestration tools (Airflow, Prefect, dbt, or similar)
- Comfort with scripting and lightweight development (Python, SQL, or similar) for pipeline development and data quality tooling
- Familiarity with cloud platforms (AWS, Azure, GCP) and experience evaluating on-premises versus cloud tradeoffs for data infrastructure
- Experience working in a controlled information environment, including familiarity with handling requirements for CUI or export-controlled technical data under ITAR or EAR
- Experience in manufacturing, industrial, or operations-heavy environments
Location
Onsite in Durham, NC. The role begins in Durham and is expected to move to Benson, NC upon completion of a new facility.