EngineerJobs.io
← Back to all jobs

Job Description

Precise Software Solutions Incorporated is seeking an AI/ML Data Engineer to help build and govern the data foundation that powers secure, AI-enabled applications for the U.S. Government. In this onsite role in Washington, DC, you will own AI-ready datasets end-to-end, from ingestion and quality controls to privacy protection, lineage, and retrieval-quality measurement, while supporting measurable and operationally secure AI.

Responsibilities

  • Design and maintain data ingestion, transformation, and processing pipelines (ETL/ELT) for AI training, evaluation, retrieval, and operations, including data migration and cleansing support.
  • Curate, validate, and version datasets; maintain dataset inventories, metadata, lineage, provenance, and ingestion logs.
  • Implement automated data-quality checks for duplication, schema changes, completeness, and freshness; maintain dataset quality scorecards and drift reports.
  • Design data models, vector stores, and embedding schemas for Retrieval-Augmented Generation (RAG) knowledge bases, including re-indexing content when sources change.
  • Measure retrieval and model quality against established baselines using metrics such as precision/recall, MRR, NDCG, and context relevance.
  • Prepare data-related deliverables, including AI model cards, ML and AI pipeline documentation, RAG/AI Pipeline Evaluation Reports, data dictionaries and embedding schema documentation, and responses to Government data calls.
  • Build secure structured-data access for AI applications (for example, natural-language-to-SQL with query validation and role-based authorization), and support dashboards and operational analytics.
  • Monitor data and ML pipelines, troubleshoot failures, and support root-cause analysis while keeping data in FedRAMP-authorized cloud regions with FIPS-validated encryption.
  • Support Responsible AI practices by preparing representative evaluation datasets, testing AI outputs for bias, accuracy, and hallucination, and documenting results to meet federal AI governance requirements.
  • Secure the AI data path, from source datasets and embeddings to prompts and logs, and support AI risk testing such as data poisoning.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Mathematics, Information Systems, or a related field, plus 5+ years of relevant experience (equivalent experience may substitute for the degree).
  • 5+ years of hands-on data engineering or database development experience, including data modeling, SQL, and ETL/ELT pipeline development.
  • Strong proficiency in Python; experience with data processing frameworks such as pandas and Spark and workflow orchestration tools such as Airflow.
  • 1+ year of experience preparing data for Generative AI or machine learning, including embeddings, vector databases, RAG knowledge bases, or training and evaluation datasets.
  • Experience implementing data quality, lineage, metadata management, and data governance controls.
  • Experience protecting sensitive data (PII), including masking, minimization, and access controls.
  • Experience with a major cloud data platform (AWS, Azure, or Google Cloud), Git, and CI/CD tools.
  • Must be a U.S. citizen or lawful permanent resident (green card holder).
  • Must reside in the Washington, DC metropolitan area and be able to work on-site at Government offices.
  • Must be able to obtain and maintain a Public Trust background investigation.

Technologies

  • Python, pandas, Spark, Airflow
  • AWS, Azure, Google Cloud
  • Git, CI/CD
  • Vector stores, vector databases, pgvector, OpenSearch
  • FIPS-validated encryption, FedRAMP-authorized cloud regions
  • Natural-language-to-SQL

Benefits

  • Comprehensive Health Benefits (Medical, Dental and Vision)
  • Flexible Spending Accounts (FSA) & Health Savings Account (HSA)
  • Retirement Plan with 4% match and discretionary match at year end
  • Paid Time Off (PTO): 15 days of PTO accrued per year; 7 holidays + 3 Floating holidays; 2 Innovation days (paid training days)
  • Short Term and Long-Term Disability
  • Paid Parental Leave
  • Paid Jury Duty leave
  • Life and AD&D Insurance
  • Critical Illness Insurance
  • Training and Development
  • Wellness Incentives & Discount programs
  • Employee Referral Program
  • Annual Charity Donation Match
  • Awards and Recognition

Preferred Qualifications

  • Master’s degree in a related field and 7+ years of data engineering experience, including support of federal agency programs.
  • Experience evaluating retrieval quality and building RAG pipelines with vector stores (for example, pgvector, OpenSearch).
  • Experience with MLOps tooling, including model and dataset versioning, feature stores, or model registries.
  • Generative AI, LLM security, or data certification (for example, Databricks Generative AI Engineer, Snowflake SnowPro, AWS, Microsoft Azure, or Google Cloud).
  • An active Public Trust or prior federal background investigation.

Location: Washington, DC (onsite)
Experience: 5+ years
Education: Bachelor’s degree

Similar Jobs