EngineerJobs.io
← Back to all jobs

Job Description

Swingtech is supporting the U.S. Department of Labor with AI/ML solutions that rely on reliable, secure, and well-governed data. As an AI/ML Data Engineer in a hybrid role based in Washington, DC, you will develop and sustain the data foundations that enable analytics, document intelligence, RAG, and production AI capabilities.

This position focuses on structured, semi-structured, and unstructured data pipelines, including document ingestion workflows, data modeling and governance controls, and operational support for secure training, validation, evaluation, and production datasets.

Responsibilities

  • Design, build, test, deploy, and maintain scalable data pipelines for batch, streaming, near-real-time, and event-driven workloads.
  • Integrate approved agency data sources, APIs, file stores, document repositories, relational databases, data lakes, data warehouses, and authorized external sources.
  • Develop ETL/ELT pipelines for extraction, validation, transformation, normalization, enrichment, de-identification, metadata management, and loading.
  • Implement document-ingestion pipelines supporting OCR, parsing, classification, metadata extraction, PII detection/redaction, chunking, embeddings, vector indexing, and retrieval workflows.
  • Create and maintain data models, schemas, data dictionaries, metadata structures, catalog records, and data-quality controls.
  • Implement data lineage, source provenance, dataset versioning, retention, access controls, and auditability across training, validation, evaluation, and production datasets.
  • Preserve separation of training, validation, and final evaluation datasets through controlled access, versioning, and documented lifecycle processes.
  • Develop and monitor data-quality measures including completeness, accuracy, timeliness, duplication, validity, freshness, distribution drift, and labeling quality.
  • Apply data minimization, masking, encryption, access controls, de-identification, and least-privilege safeguards for PII, CUI, and other protected DOL data.
  • Collaborate with AI/ML Engineers to improve retrieval quality, embeddings, vector stores, hybrid search, reranking, citation traceability, and knowledge-base refresh processes.
  • Develop data-pipeline runbooks, technical documentation, source inventories, lineage artifacts, data-quality reports, and operational support procedures.
  • Support security, privacy, ATO, Responsible AI, incident response, MLOps, monitoring, and release-readiness activities.

Requirements

  • Bachelor’s degree in computer science, data engineering, data science, information systems, software engineering, mathematics, or a related technical discipline.
  • At least 4 years of experience in data engineering, database development, analytics engineering, ETL/ELT development, data-platform implementation, or related work.
  • Strong SQL and Python development skills.
  • Experience designing data pipelines and integrating APIs, databases, file systems, cloud storage, data warehouses, or data lakes.
  • Experience with data modeling, metadata, data quality, data lineage, data transformation, monitoring, and operational support.
  • Familiarity with AWS, Azure, Google Cloud, or equivalent cloud data services.
  • Knowledge of secure data-handling practices, including access control, encryption, data masking, PII protection, and logging.
  • Willingness to work 3 days onsite at a customer site in Washington, DC.

Technologies

  • SQL, Python
  • AWS, Azure, Google Cloud
  • AWS Glue, S3, Athena, Redshift, Lake Formation
  • Azure Data Factory, Azure Data Lake Storage
  • Databricks, Snowflake, BigQuery
  • Vector databases: OpenSearch, pgvector, Pinecone, Weaviate, Milvus, Chroma, FAISS
  • RAG, OCR
  • Security and governance: FISMA, FedRAMP, NIST 800-53, NIST 800-171

Benefits

  • 15 PTO days
  • 11 paid holidays
  • Medical Insurance with 3 options (HSA with $600 employer contribution)
  • Dental Insurance with no age limit orthodonture
  • Vision Insurance through EyeMed in and out of network coverage
  • Short Term and Long-Term Disability coverage with 100% premium support
  • Life insurance and AD&D with 100% premium support
  • Supplemental Life Insurance
  • Critical Care and Accident Insurance availability
  • Pet Insurance through Nationwide
  • Employee Assistance Program
  • 401k with enrollment from day one. 4% deferral by company
  • $1500 Annual Training Budget
  • $1500 Referral bonus
  • Eligibility for annual merit and discretionary bonus
  • Flexible work arrangements

Preferred Qualifications

  • Experience with AWS Glue, S3, Athena, Redshift, Lake Formation, Azure Data Factory, Azure Data Lake Storage, Databricks, Snowflake, BigQuery, or equivalent platforms.
  • Experience with vector databases or vector-search capabilities, including OpenSearch, pgvector, Pinecone, Weaviate, Milvus, Chroma, FAISS, or similar tools.
  • Experience with RAG, document intelligence, OCR, enterprise search, knowledge management, document classification, or content-ingestion pipelines.
  • Familiarity with Federal data governance, FedRAMP, FISMA, NIST 800-53, NIST 800-171, CUI, Privacy Act, and records-management requirements.

Similar Jobs