Senior Machine Learning Engineer
Job Description
Senior Machine Learning Engineer to lead LLM powered application development on AWS, design production grade ML and LLM services, and own the end-to-end model lifecycle with MLOps.
Responsibilities
- Develop LLM applications by building RAG pipelines, prompt orchestration, tools and agents, safety rails, and evaluation harnesses; track latency, cost, and quality.
- Own the ML lifecycle from data curation and feature engineering through training and fine-tuning (LoRA/QLoRA), A/B testing, deployment, monitoring, and continuous improvement of models and prompts.
- Productionize on AWS by delivering scalable services on EKS, ECS, and Lambda; leverage SageMaker, Bedrock, EMR, MSK, and Step Functions; implement observability and cost controls.
- Establish MLOps and governance practices, including CI/CD for models (MLflow/Kedro/SageMaker Pipelines), model/version registries, data and prompt lineage, evaluation gates, and responsible AI controls.
- Collaborate across Brightly to translate asset-management use cases into ML/LLM solutions; partner with product managers and UX to ship customer-visible features that enhance reliability, safety, and sustainability.
- Conduct Exploratory Data Analysis on structured, semi-structured, and unstructured datasets to uncover patterns, correlations, feature importance, and data quality issues.
- Perform deep research on asset-related and domain-specific datasets to understand root causes, trends, and predictive signals.
- Maintain a pragmatic, product-oriented approach with a bias toward measurable outcomes and rapid stakeholder-aligned iteration.
- Demonstrate engineering excellence by writing production-grade Python, designing reliable APIs/services, and upholding testing and observability standards.
- Provide collaborative leadership by mentoring peers and influencing architecture across teams.
Requirements
- 8–10 years of total software or ML engineering experience, with at least 2 years building and operating ML systems in production.
- 1+ years hands-on LLM application development (RAG, fine-tuning, prompt engineering, evaluators/guards, agentic workflows) using Langchain and Langgraph.
- AWS proficiency (3+ years): strong with core services (EKS/ECS, Lambda, S3, DynamoDB or RDS, Step Functions, IAM) and ML stack (SageMaker, Bedrock or HF on AWS).
- Modeling and frameworks: Python, PyTorch, Hugging Face ecosystem; vector stores (OpenSearch, PGVector, Pinecone), embeddings, retrieval, and evaluation metrics for NLP/LLMs.
- MLOps: CI/CD for ML, model registries, experiment tracking, telemetry/monitoring, automated retraining; Docker/Kubernetes; GitHub Actions or GitLab CI.
- Data engineering fluency: ETL/ELT, streaming/batch (Spark/Flink), data quality and governance controls for ML.
Technologies
- Python, PyTorch, Hugging Face ecosystem, Langchain, Langgraph
- SageMaker, Bedrock
- OpenSearch, PGVector, Pinecone
- Spark, Flink
- AWS core services: EKS, ECS, Lambda, S3, DynamoDB, RDS, Step Functions, IAM
- EMR, MSK
- CloudWatch, OpenTelemetry
- Docker, Kubernetes
- GitHub Actions, GitLab CI
- MLflow, Kedro, SageMaker Pipelines
NICE TO HAVE
- Experience with distributed training (FSDP, DeepSpeed), RLHF, or optimization on Inferentia/Trainium.
- Exposure to sustainability or asset management and intelligent operations domains.
- Familiarity with security and compliance considerations for enterprise ML systems.