Pyspark/Python Data Engineer
Job Description
At Tata Consultancy Services, we are seeking a PySpark and Python Data Engineer for an onsite role in Irving, TX. This position offers a competitive annual salary of USD 100,000 to 120,000 and the opportunity to design, build, and optimize scalable PySpark data pipelines. You will work to ensure data quality, integrity, and performance while collaborating with Data Analysts, Architects, and DevOps to deploy reliable data applications and contribute to CI/CD workflows.
What we offer
Onsite in Irving, Texas with a salary range of USD 100,000–120,000 per year. You will engage with a modern tech stack that includes PySpark, Spark SQL, DataFrame API, Python, and SQL, plus cloud and data tools such as Amazon S3, AWS Glue, AWS EMR, Airflow, Snowflake and Git. The role emphasizes end-to-end data pipeline design, cross-functional collaboration, and the chance to tune performance, reliability, and data quality in production environments.
Responsibilities
- Design, develop, and maintain ETL and ELT pipelines using PySpark to ensure scalable data processing.
- Implement optimized PySpark transformations using DataFrames and Spark SQL to maximize throughput.
- Develop reusable Python-based data processing components that can be leveraged across projects.
- Ensure data quality, integrity, and performance across all pipelines.
- Debug, profile, and optimize PySpark jobs to meet production requirements.
- Collaborate with Data Analysts, Architects, and DevOps to deliver robust data solutions.
- Contribute to CI/CD pipelines and deployment workflows for data applications.
- Monitor and troubleshoot data workloads in production environments to maintain reliability.
Requirements
- Strong hands-on experience with PySpark, including Spark SQL and the DataFrame API.
- Advanced proficiency in Python for data processing, performance tuning, and modular coding.
- Solid understanding of ETL design patterns and data pipeline architecture.
- Working knowledge of SQL for data transformation and analysis.
- Experience with data processing in distributed environments.
- 3-8 years of experience in Data Engineering or PySpark development.
- Proven hands-on project experience in PySpark and Python.
- Bachelor of Computer Science degree.
Technologies
- PySpark
- Spark SQL
- DataFrame API
- Python
- SQL
- Amazon S3
- AWS Glue
- AWS EMR
- Airflow
- Snowflake
- Git