Principal Machine Learning Engineer
Ai Ml
Artificial Intelligence
Data Pipeline
Data Platform
Data Processing
DevOps
Engineering
Generative AI
Large Language Models
Machine Learning
Machine Learning Engineer
Ml Ops
MLOps
Model Monitoring
Model Optimization
Model Serving
Performance Engineering
Platform Engineering
Programming
Software Engineering
Training Pipelines
Job Description
The Principal Machine Learning Engineer will own the end-to-end machine learning infrastructure behind real-time compliance enforcement systems, including training, evaluation, and production serving. This is a hybrid role based in the New York City Metro area, with 3 days onsite.
Role Focus
Own ML infrastructure for real-time compliance enforcement systems, driving model training, evaluation, and production serving. The work includes building reliable pipelines and deployment mechanisms that support safe model updates and ongoing quality monitoring.
Responsibilities
- Build and own training pipelines covering data preparation, reproducible fine-tuning runs, experiment tracking, and release automation
- Build evaluation infrastructure with automated evaluation runs, regression gates, dashboards, and dataset versioning
- Own production model serving including low-latency inference, batching, optimization, autoscaling, and cost management
- Ship model updates safely using versioning, canarying, rollback, and drift monitoring
- Create repeatable adaptation workflows to update models for new domains and customer needs
- Turn expert labels and reviewer feedback into clean training and evaluation datasets
- Set the engineering bar for ML infrastructure as the team grows
Requirements
- 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production
- Hands-on experience with the modern LLM stack: PyTorch, distributed training, fine-tuning at scale (for example LoRA, SFT), and inference engines such as vLLM or TensorRT-LLM
- Experience building evaluation harnesses, regression gates, or dataset pipelines, with strong understanding of precision, recall, and calibration
- Proven ownership of production model serving with real constraints for latency, reliability, and cost
- Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability
- Ability to scope work, ship frequently, and make pragmatic build-vs-buy decisions
- Experience collaborating closely with research partners and defining clear interfaces
Technologies
- PyTorch
- LoRA
- SFT
- vLLM
- TensorRT-LLM
- Python
- Containers
- CI/CD
- Cloud infrastructure
- Observability
Additional Skills That May Help
- Experience productionizing small or specialized language models
- Experience with structured-output serving or constrained decoding in production
- Prior work in regulated or high-stakes domains (fintech, healthcare, legal, trust and safety)
- Experience deploying models into customer-controlled environments
Compensation and Location
- Pay range: USD 200,000 - 250,000 per year
- Location: New York, NY (hybrid, 3 days onsite)
- Work location details: Hybrid remote in New York, NY 10001
How to Apply
This role may fill quickly. Submit your resume to be considered.