Data Engineer
Job Description
Data Engineer role at Todata Analytics in Omaha, Nebraska, focused on governed ingestion and transformation pipelines for products and AI in a regulated environment.
Responsibilities
- Build and maintain data ingestion and transformation pipelines across the platform
- Design and evolve data models that support products and analytics
- Implement data governance and access controls suitable for regulated use cases, with correct client data isolation
- Apply engineering fundamentals including CI/CD, consistent naming conventions, and environment separation, and enforce them across the codebase
- Work with product and software teams to translate client requirements into well-modeled, reusable datasets
- Support infrastructure that makes governed data available to downstream consumers, including AI products
- Monitor data quality and data lineage, identify issues early, and respond before problems impact clients
Requirements
- 3+ years in data engineering, including hands-on Databricks experience (Spark, Delta Lake, Unity Catalog)
- Strong SQL skills: complex queries, joins, window functions, query performance optimization, and dimensional/star-schema modeling from ambiguous inputs
- Solid Python experience for data processing (examples include PySpark and pandas) and scripting
- Hands-on Databricks experience with notebooks, Delta Lake, jobs, and clusters
- Excellent debugging and problem-solving, with the ability to trace root causes through logs, code, and data
- Proven commitment to data governance and multi-tenant isolation, including experience explaining how datasets are kept separated and permissioned
- Experience with CI/CD, version control, and disciplined naming and environment conventions in a data context
- Track record of independently delivering foundational infrastructure work without close supervision
- Clear written communication, including documenting builds and changes
Technologies
- SQL
- Python
- PySpark
- pandas
- Databricks
- Spark
- Delta Lake
- Unity Catalog
- CI/CD
Nice to Have
- Experience in a HIPAA, SOC 2, or other regulated data environment
- Familiarity with healthcare/clinical research or financial/accounting data domains
- Exposure to enabling AI/LLM consumers of a governed semantic layer
- Databricks certification
Compensation & Location
- Location: Omaha, NE (onsite)
- Salary: USD 80,000 - 140,000 per year
- Minimum experience: 3 years
Benefits
- 401(k) matching
- Dental insurance
- Health insurance
- Life insurance
- Paid time off
- Professional development assistance
- Retirement plan
- Vision insurance