Data Engineer
Job Description
Strategic Innovation Group LLC is hiring a Senior Data Engineer to help build the data core of a new analytics platform for a federal government client. This onsite role in Arlington, VA focuses on delivering an on-premises PostgreSQL data layer in a FedRAMP Moderate environment, with strong emphasis on data correctness, reusable ingestion, and operations that client teams can maintain.
Working with the data governance lead and security lead, you will design a governed schema and shared data vocabulary, implement ingestion pipelines that normalize and validate heterogeneous inputs, and ensure analytic rules are deterministic and inspectable. The result: dashboards, analytic rules, and reporting backed by data you can trace, govern, and keep current.
Responsibilities
- Design and implement a centralized PostgreSQL data layer with a governed schema and a shared data vocabulary defined with the data governance lead.
- Architect a scalable ingestion pipeline in Python using orchestration frameworks such as Airflow, Dagster, Prefect, or comparable tools to collect, normalize, validate, and persist structured and unstructured data from agency financial-management, program-management, and payment systems, plus Government-wide sources.
- Build configuration-driven connectors (REST APIs, flat files, database replication, SFTP) reusable across agencies without source-code changes, and deploy through the program CI/CD pipeline (GitLab, Jenkins or Azure DevOps) with automated build, test, and security scanning to the client’s Rancher-managed Kubernetes platform.
- Implement PII anonymization, data-handling, and access controls at ingestion in coordination with the security lead, aligned to NIST SP 800-53 Moderate controls.
- Implement defined analytic rules and threshold-based indicators as deterministic, inspectable SQL that runs on a scheduled refresh, and document the data and criteria behind each rule.
- Lead an error-detection methodology for data-quality issues by identifying, categorizing, tracking, resolving, and reporting defects and anomalies across the lifecycle.
- Capture metadata and data lineage, and expose data-currency indicators (source, last refreshed, next refresh) to the user interface.
- Set coding, code-review, testing, and documentation standards for the data team; review the Data Engineer’s work; and collaborate with application developers on API access to the data core.
- Troubleshoot and resolve pipeline failures, data discrepancies, and query performance bottlenecks.
- Produce technical documentation, data dictionaries, and runbooks, delivering knowledge transfer so client staff can operate, maintain, and extend pipelines with minimal contractor reliance.
- Participate in Agile ceremonies, sprint planning, code reviews, demonstrations, and continuous improvement activities.
Requirements
- Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field; equivalent experience may be considered.
- 8+ years of professional data engineering experience, including 3+ years as a technical lead or senior engineer on a production data platform (or 6+ years with a Master’s degree).
- Expert-level SQL and PostgreSQL data modeling, including normalized and analytic schemas, partitioning, indexing, and performance tuning.
- Strong Python for pipeline development and hands-on experience with an orchestration framework such as Airflow, Dagster, or Prefect.
- Demonstrated experience designing reusable, configuration-driven ingestion from heterogeneous sources with schema validation and data-quality checks.
- Experience deploying pipelines in containerized environments (e.g., Docker, Kubernetes) with Git-based version control and CI/CD.
- Experience handling PII and sensitive financial or program data under NIST SP 800-53 or equivalent controls.
- Experience working in Agile software development environments and collaborating with technical and non-technical stakeholders in a federal contracting environment.
- Strong written communication skills for data-model documentation, runbooks, and knowledge transfer.
Technologies
PostgreSQL, Python, Airflow, Dagster, Prefect, REST APIs, SFTP, GitLab, Jenkins, Azure DevOps, Rancher, Kubernetes, Docker, CI/CD, Git-based version control, NIST SP 800-53, SQL
Benefits
- Great work/life balance
- Eligibility for performance-based participation in cash bonuses
- Potential to participate in growth of the company through incentives
- Excellent benefits, including health, dental, vision, generous PTO, a 401(k) with match, life insurance, short- and long-term disability, and a health savings account (HSA)
Preferred Qualifications
- Experience supporting federal civilian, financial-oversight, budget, or program-management IT systems.
- Experience integrating data from enterprise financial systems (e.g., SAP S/4HANA, Oracle, or comparable ERP) and Treasury or payment platforms.
- Experience with grants management or federal financial systems (e.g., GrantSolutions, eRA, Treasury payment systems) highly preferred.
- Experience extracting structured data from unstructured documents (PDF, DOCX) at scale.
- Experience with metadata management, data-lineage, and data-catalog tooling.
- Experience deploying to Kubernetes (Rancher preferred) in an on-premises federal environment with restricted internet access.
- Familiarity with Splunk log integration and NIST SP 800-53 audit-logging requirements.