Vantage Data Centers Management Company LLC is seeking a Sr Data Engineer, Data Analytics & Intelligence, NA to build, operate, and scale governed data foundations that support Operations across North America. In this role, you will design Azure-based data pipelines and curated datasets, including semantic-model inputs, enabling enterprise reporting, operational intelligence, and AI-enabled use cases.
This position will be based onsite in Denver, CO in alignment with the flexible work policy, with 3 days on site required and 2 days flexible. Salary for the role is USD 130,000 - 155,000 per year, based on Colorado market data and subject to variation based on qualifications and experience.
What you will do
- Design, build, and maintain reliable, scalable data pipelines using Python and PySpark on the Microsoft Azure data platform.
- Develop and operate batch and incremental pipelines using Azure Data Factory, with Azure Data Lake Storage Gen2 as the primary data store.
- Build and maintain curated lakehouse / gold-layer datasets and semantic-model inputs for governed operational insights and AI-enabled consumption.
- Implement SQL- and Spark-based transformations to produce curated datasets supporting enterprise reporting, analytics, and downstream AI-enabled insight preparation.
- Own assigned pipelines and datasets, including monitoring, troubleshooting, performance optimization, documentation, and production support.
- Work with Azure Synapse and Microsoft Fabric / Lakehouse patterns where applicable, along with related Azure analytics services.
- Prepare structured operational data for AI-enabled use cases by documenting business rules, source lineage, data reliability constraints, known quality limitations, and data dictionary definitions.
- Support source visibility and confidence context, including Data Reliability & Trust Indicator integration where applicable.
- Contribute to ontology, taxonomy, semantic model, and data dictionary alignment across operational context, KPIs, incidents, work orders, and other enterprise domains.
- Collaborate with business analysts, operations SMEs, data stewards, IT Global, and cross-functional stakeholders to translate requirements into working data solutions.
- Apply data governance, security, access control, data classification, and engineering standards to ensure scalable, compliant solutions.
- Identify, document, and route data-quality issues to accountable owners to improve source correction rather than hiding defects downstream.
- Participate in code reviews, technical discussions, sprint planning, and platform improvement initiatives.
- Develop and maintain PySpark notebooks and jobs for ingesting, transforming, validating, and curating data.
- Create and modify Azure Data Factory pipelines for batch and incremental ingestion.
- Implement Spark transformations to write curated outputs to Azure Data Lake Storage Gen2 and/or Fabric Lakehouse patterns using established structures, naming conventions, and governance.
- Create and maintain SQL views, tables, lakehouse objects, and semantic-model inputs for analytics and operational intelligence.
- Prepare datasets for Fabric Data Agent / AI agent use cases by documenting business rules, joins, grain, quality limitations, source lineage, and operational definitions.
- Respond to pipeline failures, data validation issues, operational alerts, and data-quality escalations with root-cause analysis and remediation steps.
- Perform performance tuning of Spark jobs and SQL workloads (partitioning, filtering, incremental logic, query optimization, and resource-aware design).
- Validate data outputs with business partners and operations SMEs, and address discrepancies through documented correction paths.
- Support observability, logging, and auditability practices for data pipelines and AI-consumable datasets where applicable.
- Commit code using Git, follow branching standards, participate in pull request reviews, and support CI/CD with GitHub, Azure DevOps, or similar tools.
- Maintain documentation for pipelines, datasets, data contracts, data dictionaries, business rules, and operational runbooks as changes are made.
- Execute sprint backlog items, raising risks, dependencies, or blockers early.
- Complete additional duties as assigned by management.
Core requirements
- Bachelorβs degree in Engineering, Computer Science, Data Analytics, or a related field, or equivalent experience.
- Minimum 5β8 years of experience in data engineering, analytics engineering, or a closely related technical data role.
- Proficiency in Python and PySpark for data pipelines and transformations.
- Proficiency in SQL for querying, transformations, model validation, and data quality checks.
- Solid understanding of ETL/ELT, transformation patterns, data integration, incremental processing, and production support.
- Experience analyzing enterprise data sources to identify relationships, transformations, business rules, grain, ownership, and quality constraints.
- Experience building on Microsoft Azure, including Azure Data Factory, Azure Synapse, Azure Data Lake Storage Gen2, and Microsoft Fabric / Lakehouse patterns.
- Working knowledge of data modeling fundamentals, including fact/dimension tables and semantic models.
- Experience supporting governed data products, including metadata, lineage, issue documentation, access-control awareness, validation, and operational runbooks.
- Experience with source control and CI/CD using GitHub or Azure DevOps.
- Strong communication skills and the ability to collaborate with IT, business SMEs, and data governance partners in a fast-paced environment.
- Experience working in Agile environments and using tools such as Jira or similar.
- Travel is expected up to 10%, and may increase over time.
Technologies
- Python, PySpark, Microsoft Azure
- Azure Data Factory, Azure Data Lake Storage Gen2, Azure Synapse
- Microsoft Fabric, Fabric Lakehouse, SQL, Spark
- Git, GitHub, Azure DevOps, Jira, CI/CD
- Microsoft Fabric / Lakehouse patterns
Benefits
- Medical, dental, and vision coverage
- Life and AD&D
- Short and long-term disability coverage
- Paid time off
- Employee assistance
- Participation in a 401k program that includes company match
- Above market total compensation package
- Comprehensive suite of health and welfare, retirement, and paid leave benefits
Salary range
$130k - $155k (Colorado market-based; may vary by location). Compensation depends on qualifications, skills, competencies, and experience and may fall outside the range shown.
Additional details
- Reasonable accommodations may be made for individuals with disabilities to perform essential functions.
- Occasionally required to stand, walk, sit, use hands to handle or feel objects, reach, climb stairs, balance, stoop or kneel, talk and hear.
- Occasionally required to lift and/or move up to 25 pounds.
- This position is eligible for company benefits including medical, dental, and vision, life and AD&D, short and long-term disability, paid time off, employee assistance, 401k with company match, and other voluntary benefits.
Desired qualifications
- Experience with distributed data processing frameworks including Apache Spark.
- Experience preparing governed data products for AI-enabled use cases, including Microsoft Fabric Lakehouse, semantic models, Fabric/Data Agent patterns, and ontology or taxonomy alignment for explainable AI outputs.
- Familiarity with data observability, metadata management, lineage, data contracts, reliability indicators, and operational best practices in production environments.
- Familiarity with additional Azure services such as Azure Functions or Logic Apps.
- Experience supporting data platform enhancement, refactoring, modernization, or reusable architecture initiatives.
- Experience working with structured and unstructured operational sources such as enterprise applications, operational workflows, documents, dashboards, and knowledge assets.
- Experience working in a scaling or fast-paced organization where priorities evolve quickly and practical delivery discipline is required.