EngineerJobs.io
← Back to all jobs

Job Description

Placement Force is seeking a Senior Machine Learning Engineer to own the complete lifecycle of a production large language model. The role blends hands-on model training, scalable deployment, and thorough evaluation to ensure the shipped system meets real-world objectives. This is a remote position that demands a proven ability to deliver robust ML systems in production environments.

Location: Remote

Compensation: USD 160,000 - 260,000 per year

Minimum experience: 4 years

Responsibilities

  • Architect and run pretraining, continued pretraining, and fine-tuning experiments, including SFT, DPO/RLHF-style alignment, LoRA/QLoRA, and full-parameter methods, all tied to clear, measurable objectives
  • Own data pipeline decisions that meaningfully impact model quality, covering data curation, de-duplication, mixture weighting, and contamination checks against evaluation sets
  • Conduct and interpret distributed training across multi-GPU and multi-node setups using frameworks such as FSDP, DeepSpeed, or Megatron-style parallelism; diagnose failures that appear only at scale
  • Make and defend practical tradeoffs between model size, training cost, and downstream performance
  • Take trained models to production with quantization, batching strategies, KV-cache management, and serving framework choices (for example vLLM, TensorRT-LLM, or TGI) while meeting explicit latency, throughput, and cost targets
  • Design for LLM serving failure modes, including tail latency under load, graceful degradation, prompt injection surfaces, and safe fallback behavior
  • Establish and maintain production-grade operational practices such as monitoring, alerting, and rollback procedures, treating the model as a critical service rather than a one-off export
  • Develop evaluation harnesses that go beyond published benchmarks, incorporating task-specific evaluations aligned with real product use cases
  • Design human evaluation protocols to complement automated metrics and determine when each approach is appropriate
  • Own regression detection to catch quality drops caused by new checkpoints, prompt template changes, or serving optimizations before users notice
  • Contribute to safety and robustness assessments, including measuring hallucination rates, adversarial testing, and behavior under distribution shift as an integral part of the release process

Requirements

  • At least 4 years of applied ML engineering experience, with a minimum of 2 years working directly on large language models in production
  • Hands-on production deployment experience, including shipping a model that served live traffic and being able to discuss latency, cost, and quality tradeoffs
  • Strong software engineering fundamentals
  • Fluency with the modern LLM tooling landscape, including training frameworks, serving frameworks, and evaluation tooling
  • Comfort with ambiguity and the ability to define what constitutes “good” model behavior in areas lacking established benchmarks

Technologies

  • FSDP
  • DeepSpeed
  • Megatron-style parallelism
  • vLLM
  • TensorRT-LLM
  • TGI

Similar Jobs