AI Engineer
Job Description
LawPro.ai is building AI-powered data insights and analytics, and they’re hiring an AI Engineer in Virginia (onsite) to own the evaluation, selection, and continuous optimization of large language models and AI processes. This role blends AI research with production engineering, with end-to-end responsibility for model transitions, pipeline optimizations, monitoring, and documentation.
What you’ll deliver
- Run a systematic, ongoing process to evaluate new and emerging LLMs on accuracy, relevancy, speed, and cost, continuously benchmarking against the specific tasks in the orchestration pipeline.
- Design and operate rigorous Evals and build an internal EvalOps culture to measure output accuracy, relevance, faithfulness, and speed, with focused effort on reducing hallucinations in medical record summarization and legal document analysis.
- Track provider changes across the LLM landscape to identify deprecation timelines, then execute full transitions by integrating replacement models into the production pipeline and updating for model behavior differences.
- Improve LLM-based orchestration for document understanding, medical record summarization, case chronology generation, and drafting support by implementing code changes, deployments, and production validation from start to finish with a bias for surgical execution.
- Communicate evaluation findings with product and GTM stakeholders, then lead the technical implementation directly to ensure frictionless handoffs between discovery, staging, and live production deployments.
- Ship model changes into production by writing integration code, managing deployments, running validation tests, and ensuring clean rollouts.
- Implement monitoring and observability for model performance, including benchmarking outputs and cost, detecting drift, and reporting continuously to management using micro-benchmarking to track token-level latency, output drift, and cost efficiency across pipeline components.
- Maintain thorough documentation covering evaluation methodologies, model comparison results, transition decisions, and runbooks for the systems you own.
What you bring
- 5+ years of AI/ML engineering experience evaluating, fine-tuning, and deploying large language models in production, including building and deploying models to AWS or GCP infrastructure at scale.
- Hands-on development and implementation of multiple RAG solutions.
- Hands-on experience leveraging embedding models and vector databases.
- Hands-on experience building agentic workflows and implementing practical EvalOps or Evals-as-a-Service architecture.
- Deep familiarity with the LLM ecosystem, with ability to assess model capabilities, limitations, and fit using tradeoffs across routing, cost, quality, speed, and capability.
- Proven experience designing evaluation frameworks for output quality measurement, including hallucination detection in high-stakes domains (legal, medical, or similar).
- Strong software engineering foundation with production experience building LLM orchestration frameworks and multi-model pipelines.
- Comfort working in a fast-paced, high-ambiguity environment with strong ownership, tight feedback loops, and a bias toward systematic process-building.
- Excellent communication skills to translate complex evaluation findings into clear recommendations for engineering, product, and non-technical stakeholders.
- Bonus: experience with unstructured medical or legal document processing, or background in classical ML (statistics, embeddings, retrieval-augmented generation).
Technology focus
AWS, GCP