Senior Machine Learning Engineer
Job Description
Placement Force is seeking a Senior Machine Learning Engineer to own the complete lifecycle of a production large language model. The role blends hands-on model training, scalable deployment, and thorough evaluation to ensure the shipped system meets real-world objectives. This is a remote position that demands a proven ability to deliver robust ML systems in production environments.
Location: Remote
Compensation: USD 160,000 - 260,000 per year
Minimum experience: 4 years
Responsibilities
- Architect and run pretraining, continued pretraining, and fine-tuning experiments, including SFT, DPO/RLHF-style alignment, LoRA/QLoRA, and full-parameter methods, all tied to clear, measurable objectives
- Own data pipeline decisions that meaningfully impact model quality, covering data curation, de-duplication, mixture weighting, and contamination checks against evaluation sets
- Conduct and interpret distributed training across multi-GPU and multi-node setups using frameworks such as FSDP, DeepSpeed, or Megatron-style parallelism; diagnose failures that appear only at scale
- Make and defend practical tradeoffs between model size, training cost, and downstream performance
- Take trained models to production with quantization, batching strategies, KV-cache management, and serving framework choices (for example vLLM, TensorRT-LLM, or TGI) while meeting explicit latency, throughput, and cost targets
- Design for LLM serving failure modes, including tail latency under load, graceful degradation, prompt injection surfaces, and safe fallback behavior
- Establish and maintain production-grade operational practices such as monitoring, alerting, and rollback procedures, treating the model as a critical service rather than a one-off export
- Develop evaluation harnesses that go beyond published benchmarks, incorporating task-specific evaluations aligned with real product use cases
- Design human evaluation protocols to complement automated metrics and determine when each approach is appropriate
- Own regression detection to catch quality drops caused by new checkpoints, prompt template changes, or serving optimizations before users notice
- Contribute to safety and robustness assessments, including measuring hallucination rates, adversarial testing, and behavior under distribution shift as an integral part of the release process
Requirements
- At least 4 years of applied ML engineering experience, with a minimum of 2 years working directly on large language models in production
- Hands-on production deployment experience, including shipping a model that served live traffic and being able to discuss latency, cost, and quality tradeoffs
- Strong software engineering fundamentals
- Fluency with the modern LLM tooling landscape, including training frameworks, serving frameworks, and evaluation tooling
- Comfort with ambiguity and the ability to define what constitutes “good” model behavior in areas lacking established benchmarks
Technologies
- FSDP
- DeepSpeed
- Megatron-style parallelism
- vLLM
- TensorRT-LLM
- TGI