EngineerJobs.io
← Back to all jobs

Job Description

The Principal Machine Learning Engineer will own the end-to-end machine learning infrastructure behind real-time compliance enforcement systems, including training, evaluation, and production serving. This is a hybrid role based in the New York City Metro area, with 3 days onsite.

Role Focus

Own ML infrastructure for real-time compliance enforcement systems, driving model training, evaluation, and production serving. The work includes building reliable pipelines and deployment mechanisms that support safe model updates and ongoing quality monitoring.

Responsibilities

  • Build and own training pipelines covering data preparation, reproducible fine-tuning runs, experiment tracking, and release automation
  • Build evaluation infrastructure with automated evaluation runs, regression gates, dashboards, and dataset versioning
  • Own production model serving including low-latency inference, batching, optimization, autoscaling, and cost management
  • Ship model updates safely using versioning, canarying, rollback, and drift monitoring
  • Create repeatable adaptation workflows to update models for new domains and customer needs
  • Turn expert labels and reviewer feedback into clean training and evaluation datasets
  • Set the engineering bar for ML infrastructure as the team grows

Requirements

  • 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production
  • Hands-on experience with the modern LLM stack: PyTorch, distributed training, fine-tuning at scale (for example LoRA, SFT), and inference engines such as vLLM or TensorRT-LLM
  • Experience building evaluation harnesses, regression gates, or dataset pipelines, with strong understanding of precision, recall, and calibration
  • Proven ownership of production model serving with real constraints for latency, reliability, and cost
  • Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability
  • Ability to scope work, ship frequently, and make pragmatic build-vs-buy decisions
  • Experience collaborating closely with research partners and defining clear interfaces

Technologies

  • PyTorch
  • LoRA
  • SFT
  • vLLM
  • TensorRT-LLM
  • Python
  • Containers
  • CI/CD
  • Cloud infrastructure
  • Observability

Additional Skills That May Help

  • Experience productionizing small or specialized language models
  • Experience with structured-output serving or constrained decoding in production
  • Prior work in regulated or high-stakes domains (fintech, healthcare, legal, trust and safety)
  • Experience deploying models into customer-controlled environments

Compensation and Location

  • Pay range: USD 200,000 - 250,000 per year
  • Location: New York, NY (hybrid, 3 days onsite)
  • Work location details: Hybrid remote in New York, NY 10001

How to Apply

This role may fill quickly. Submit your resume to be considered.

Similar Jobs