AI/ML Data Engineer
Job Description
Precise Software Solutions Incorporated is seeking an AI/ML Data Engineer to help build and govern the data foundation that powers secure, AI-enabled applications for the U.S. Government. In this onsite role in Washington, DC, you will own AI-ready datasets end-to-end, from ingestion and quality controls to privacy protection, lineage, and retrieval-quality measurement, while supporting measurable and operationally secure AI.
Responsibilities
- Design and maintain data ingestion, transformation, and processing pipelines (ETL/ELT) for AI training, evaluation, retrieval, and operations, including data migration and cleansing support.
- Curate, validate, and version datasets; maintain dataset inventories, metadata, lineage, provenance, and ingestion logs.
- Implement automated data-quality checks for duplication, schema changes, completeness, and freshness; maintain dataset quality scorecards and drift reports.
- Design data models, vector stores, and embedding schemas for Retrieval-Augmented Generation (RAG) knowledge bases, including re-indexing content when sources change.
- Measure retrieval and model quality against established baselines using metrics such as precision/recall, MRR, NDCG, and context relevance.
- Prepare data-related deliverables, including AI model cards, ML and AI pipeline documentation, RAG/AI Pipeline Evaluation Reports, data dictionaries and embedding schema documentation, and responses to Government data calls.
- Build secure structured-data access for AI applications (for example, natural-language-to-SQL with query validation and role-based authorization), and support dashboards and operational analytics.
- Monitor data and ML pipelines, troubleshoot failures, and support root-cause analysis while keeping data in FedRAMP-authorized cloud regions with FIPS-validated encryption.
- Support Responsible AI practices by preparing representative evaluation datasets, testing AI outputs for bias, accuracy, and hallucination, and documenting results to meet federal AI governance requirements.
- Secure the AI data path, from source datasets and embeddings to prompts and logs, and support AI risk testing such as data poisoning.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Mathematics, Information Systems, or a related field, plus 5+ years of relevant experience (equivalent experience may substitute for the degree).
- 5+ years of hands-on data engineering or database development experience, including data modeling, SQL, and ETL/ELT pipeline development.
- Strong proficiency in Python; experience with data processing frameworks such as pandas and Spark and workflow orchestration tools such as Airflow.
- 1+ year of experience preparing data for Generative AI or machine learning, including embeddings, vector databases, RAG knowledge bases, or training and evaluation datasets.
- Experience implementing data quality, lineage, metadata management, and data governance controls.
- Experience protecting sensitive data (PII), including masking, minimization, and access controls.
- Experience with a major cloud data platform (AWS, Azure, or Google Cloud), Git, and CI/CD tools.
- Must be a U.S. citizen or lawful permanent resident (green card holder).
- Must reside in the Washington, DC metropolitan area and be able to work on-site at Government offices.
- Must be able to obtain and maintain a Public Trust background investigation.
Technologies
- Python, pandas, Spark, Airflow
- AWS, Azure, Google Cloud
- Git, CI/CD
- Vector stores, vector databases, pgvector, OpenSearch
- FIPS-validated encryption, FedRAMP-authorized cloud regions
- Natural-language-to-SQL
Benefits
- Comprehensive Health Benefits (Medical, Dental and Vision)
- Flexible Spending Accounts (FSA) & Health Savings Account (HSA)
- Retirement Plan with 4% match and discretionary match at year end
- Paid Time Off (PTO): 15 days of PTO accrued per year; 7 holidays + 3 Floating holidays; 2 Innovation days (paid training days)
- Short Term and Long-Term Disability
- Paid Parental Leave
- Paid Jury Duty leave
- Life and AD&D Insurance
- Critical Illness Insurance
- Training and Development
- Wellness Incentives & Discount programs
- Employee Referral Program
- Annual Charity Donation Match
- Awards and Recognition
Preferred Qualifications
- Master’s degree in a related field and 7+ years of data engineering experience, including support of federal agency programs.
- Experience evaluating retrieval quality and building RAG pipelines with vector stores (for example, pgvector, OpenSearch).
- Experience with MLOps tooling, including model and dataset versioning, feature stores, or model registries.
- Generative AI, LLM security, or data certification (for example, Databricks Generative AI Engineer, Snowflake SnowPro, AWS, Microsoft Azure, or Google Cloud).
- An active Public Trust or prior federal background investigation.
Location: Washington, DC (onsite)
Experience: 5+ years
Education: Bachelor’s degree