Machine Learning Engineer, Frontier Data Products
Job Description
Mercor's Frontier Data Products team in New York, NY (onsite) offers the opportunity to build and own production ML systems that score, validate, and improve complex work products with imperfect labels. You will design evaluation frameworks, create feedback loops, and collaborate with backend engineers to deploy durable, long-running inference pipelines. Salary ranges from USD 130,000 to 500,000 per year.
Benefits
- Bi-annual performance bonus structure.
- Generous equity grant vested over 4 years.
- Up to $15k relocation bonus.
- $10K housing bonus if you live within 0.5 miles of our office.
- $1.5K monthly stipend for meals.
- Free Equinox membership.
- $200 monthly laundry reimbursement.
- $200 monthly personal wellness reimbursement.
- Health, Dental, Vision insurance.
Responsibilities
- Develop ML systems that score, validate, and improve intricate work products where correctness is nuanced and labels are imperfect.
- Design evaluation frameworks for tasks where ground truth is partial, delayed, or disputed.
- Build feedback loops that convert reviews, disagreements, corrections, and adjudication into measurable model and system improvements.
- Own end-to-end production ML behavior, including precision and recall tradeoffs, regression detection, drift, latency, cost, and explainability.
- Improve model quality using the right tools for the job, such as prompting, fine-tuning, retrieval, active learning, heuristics, and error analysis.
- Collaborate with backend engineers to integrate inference into durable workflows without sacrificing debuggability or human oversight.
Requirements
- Track record of shipping ML systems that improved a real product, workflow, or business metric.
- Strong instincts for model quality, evaluation design, error analysis, and production failure modes.
- Comfort operating in ambiguous problem spaces where labels are imperfect and correctness evolves.
- Good judgment about when to use prompting, fine-tuning, retrieval, human review, or a simpler product constraint.
- Solid engineering fundamentals across the full ML stack, not just modeling.
- Familiarity with LLM applications, model-assisted workflows, evaluation frameworks, or human-in-the-loop ML is a strong plus.
- Defaults to simple, inspectable ML systems that improve quickly and fail in understandable ways, not the most impressive architecture.
- Discomfort shipping a model without a clear evaluation story.
- Ability to hold ambiguity without paralysis and make reasonable bets with incomplete information.
- Focus on real-world output of the system, not just benchmarks.
Technologies
Python, Temporal, Postgres, AWS, LiteLLM
Day to Day
- Move quickly on a young, high-ownership codebase where decisions have long-term architectural weight.
- Operate across models, data, backend systems, and product surfaces; context switching is the default.
- Debug production ML failures in live, long-running workflows where silent errors matter.
- Work closely with backend engineers on a stack of Python, Temporal, Postgres, AWS, and LiteLLM.
- Balance automation confidence with human review, recognizing when to defer is as important as shipping.
What makes this role different
- The architecture is not fixed; early engineers define how quality is measured, how models and humans interact, where automation is trusted, and how the system compounds over time.
- The feedback loop is short; shipping a model behavior change directly affects what customers receive.
- You will work in a strategically central product area at Mercor at a time when frontier AI solutions to this problem are lacking.