EngineerJobs.io
← Back to all jobs

Job Description

Senior Machine Learning Engineer to lead LLM powered application development on AWS, design production grade ML and LLM services, and own the end-to-end model lifecycle with MLOps.

Responsibilities

  • Develop LLM applications by building RAG pipelines, prompt orchestration, tools and agents, safety rails, and evaluation harnesses; track latency, cost, and quality.
  • Own the ML lifecycle from data curation and feature engineering through training and fine-tuning (LoRA/QLoRA), A/B testing, deployment, monitoring, and continuous improvement of models and prompts.
  • Productionize on AWS by delivering scalable services on EKS, ECS, and Lambda; leverage SageMaker, Bedrock, EMR, MSK, and Step Functions; implement observability and cost controls.
  • Establish MLOps and governance practices, including CI/CD for models (MLflow/Kedro/SageMaker Pipelines), model/version registries, data and prompt lineage, evaluation gates, and responsible AI controls.
  • Collaborate across Brightly to translate asset-management use cases into ML/LLM solutions; partner with product managers and UX to ship customer-visible features that enhance reliability, safety, and sustainability.
  • Conduct Exploratory Data Analysis on structured, semi-structured, and unstructured datasets to uncover patterns, correlations, feature importance, and data quality issues.
  • Perform deep research on asset-related and domain-specific datasets to understand root causes, trends, and predictive signals.
  • Maintain a pragmatic, product-oriented approach with a bias toward measurable outcomes and rapid stakeholder-aligned iteration.
  • Demonstrate engineering excellence by writing production-grade Python, designing reliable APIs/services, and upholding testing and observability standards.
  • Provide collaborative leadership by mentoring peers and influencing architecture across teams.

Requirements

  • 8–10 years of total software or ML engineering experience, with at least 2 years building and operating ML systems in production.
  • 1+ years hands-on LLM application development (RAG, fine-tuning, prompt engineering, evaluators/guards, agentic workflows) using Langchain and Langgraph.
  • AWS proficiency (3+ years): strong with core services (EKS/ECS, Lambda, S3, DynamoDB or RDS, Step Functions, IAM) and ML stack (SageMaker, Bedrock or HF on AWS).
  • Modeling and frameworks: Python, PyTorch, Hugging Face ecosystem; vector stores (OpenSearch, PGVector, Pinecone), embeddings, retrieval, and evaluation metrics for NLP/LLMs.
  • MLOps: CI/CD for ML, model registries, experiment tracking, telemetry/monitoring, automated retraining; Docker/Kubernetes; GitHub Actions or GitLab CI.
  • Data engineering fluency: ETL/ELT, streaming/batch (Spark/Flink), data quality and governance controls for ML.

Technologies

  • Python, PyTorch, Hugging Face ecosystem, Langchain, Langgraph
  • SageMaker, Bedrock
  • OpenSearch, PGVector, Pinecone
  • Spark, Flink
  • AWS core services: EKS, ECS, Lambda, S3, DynamoDB, RDS, Step Functions, IAM
  • EMR, MSK
  • CloudWatch, OpenTelemetry
  • Docker, Kubernetes
  • GitHub Actions, GitLab CI
  • MLflow, Kedro, SageMaker Pipelines

NICE TO HAVE

  • Experience with distributed training (FSDP, DeepSpeed), RLHF, or optimization on Inferentia/Trainium.
  • Exposure to sustainability or asset management and intelligent operations domains.
  • Familiarity with security and compliance considerations for enterprise ML systems.

Similar Jobs