EngineerJobs.io
← Back to all jobs

Job Description

Virtues presents an onsite senior Big Data Engineer role in Irving, TX with a competitive annual salary range of $100,696.26 to $110,268.61. This position centers on designing, developing, and supporting scalable batch and real-time data pipelines across multiple data platforms. You’ll work in a collaborative environment to shape data infrastructure and drive data-driven decisions, tackling complex challenges with the Hadoop ecosystem and Spark at scale.

Responsibilities

  • Design, develop, implement, and maintain scalable, high-performance data ingestion and processing pipelines using Hadoop ecosystem technologies.
  • Develop and manage data pipelines supporting batch, real-time, streaming, and event-driven processing, including Kafka events.
  • Read data from structured, semi-structured, and unstructured sources.
  • Ingest batch data and real-time event streams, including Kafka events.
  • Perform complex data validation, cleansing, enrichment, and transformation.
  • Deliver processed data to target data stores, curated data layers, publishing zones, and downstream endpoints.
  • Develop and optimize Apache Spark applications using Scala for large-scale distributed data processing.
  • Design and implement Kafka-centric event processing and real-time data pipelines.
  • Develop streaming data transformation logic using Apache Spark Streaming and/or Spark Structured Streaming.
  • Build and maintain scalable batch processing solutions using Apache Spark.
  • Develop data processing and analytical solutions using HiveQL, Pig Latin, HBase, and custom MapReduce programs.
  • Transform and move data from raw zones to curated and published data warehouse layers.
  • Collaborate with data architects, application teams, business stakeholders, and platform teams to translate requirements into scalable solutions.
  • Work extensively with Hadoop technologies including HDFS; MapReduce; Hive; Pig; Sqoop; HBase; ZooKeeper; Oozie; Apache Spark; Scala; Flume/Flume NG; Kafka; Hue.
  • Apply knowledge of Hadoop architecture and core components such as NameNode, DataNode, HDFS, JobTracker, TaskTracker, and MapReduce programming.
  • Install, configure, integrate, and support Hadoop ecosystem components within Cloudera-based environments.
  • Ensure scalability, reliability, fault tolerance, and high performance for distributed storage and processing frameworks.
  • Monitor and optimize data pipeline performance, resource utilization, throughput, and processing efficiency.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, Data Science, or a related technical discipline.
  • 7+ years of hands-on experience with Hadoop framework and the broader Hadoop ecosystem.
  • 6+ years of hands-on experience developing data ingestion and integration solutions across multiple data platforms.
  • 5+ years of strong hands-on experience in Apache Spark with Scala-based distributed data processing.
  • 5+ years of experience in data modeling, data transformation, detailed technical design, and data integration.
  • Strong experience designing and developing large-scale batch and real-time data pipelines.
  • Strong experience with HiveQL, Pig Latin, HBase, and custom MapReduce programming.
  • Experience developing and managing Kafka-centric event-driven data pipelines.
  • Strong understanding of batch processing, stream processing, and event-driven architecture.
  • Hands-on experience with Spark Streaming and/or Spark Structured Streaming.
  • Experience installing and configuring Cloudera Hadoop ecosystem components, including Hive, HBase, ZooKeeper, Oozie, Spark, Sqoop, Flume, Pig, and Hue.
  • Strong understanding of Hadoop architecture, HDFS, distributed storage, and MapReduce concepts.
  • Strong analytical, problem-solving, debugging, and performance-tuning skills.
  • Excellent communication and collaboration skills.

Technologies

  • Hadoop
  • HDFS
  • MapReduce
  • Hive
  • Pig
  • Sqoop
  • HBase
  • ZooKeeper
  • Oozie
  • Apache Spark
  • Spark Streaming
  • Spark Structured Streaming
  • Scala
  • Flume
  • Flume NG
  • Kafka
  • Hue
  • HiveQL
  • Pig Latin
  • BigQuery
  • Cloudera
  • MapR
  • Hortonworks

Key competencies

  • Strong expertise in distributed data processing and big data architecture.
  • Deep understanding of batch, real-time, streaming, and event-driven data processing.
  • Strong hands-on programming skills in Scala and distributed data engineering frameworks.
  • Ability to design scalable, fault-tolerant, and high-performance data solutions.
  • Strong technical troubleshooting and root-cause analysis capabilities.
  • Ability to work independently while collaborating effectively with cross-functional teams.
  • Strong ownership, attention to detail, and commitment to data quality and operational excellence.

Desirable skills

  • End-to-end Hadoop administration and production support experience.
  • Hadoop infrastructure setup, software installation, configuration, upgrades, patching, monitoring, troubleshooting, and maintenance.
  • Experience administering Hadoop distributions such as Cloudera, MapR, Hortonworks.
  • Installing, configuring, and managing Hadoop ecosystem components such as Hive, Pig, HBase, ZooKeeper, Oozie, Spark, Sqoop, Flume, Hue.
  • Managing and monitoring HDFS, distributed file systems, and Hadoop clusters.
  • Managing, monitoring, scheduling, and troubleshooting MapReduce and distributed processing jobs.
  • Cluster capacity planning, resource management, health monitoring, and operational support.
  • Automating operational activities using scripting for backups, cluster monitoring, health checks, maintenance, and operational reporting.
  • Experience with version control, change management, release management, incident management, problem management, and root-cause analysis.

Compensation and location

Salary: $100,696.26 to $110,268.61 per year. Location: Irving, TX, onsite.

Similar Jobs