Java Spark Engineer

@ 2T Consulting
2T Consulting2tconsulting.com

Java Spark Engineer

Blackwells Mills, New Jersey
Posted 1 day ago

About the job

The company specializes in data infrastructure and engineering. The role involves architecting and building fault-tolerant data pipelines, leading system design, and mentoring engineers to ensure reliable and efficient big data solutions.

Requirements

  • 7+ years Java development
  • 5+ years Spark experience
  • Strong SQL and storage format skills
  • Experience with Kafka and cloud Spark
  • Knowledge of distributed systems

Qualifications

  • Bachelor’s or Master’s in CS or related
  • Expert in distributed systems
  • Proven system design experience

Full job description

Primary Responsibilities

Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java)

Lead design of batch and streaming ETL/ELT systems handling large data volumes

Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction

Set coding standards and lead code/design reviews across the team

Drive technical decisions on data architecture, storage formats, and pipeline orchestration

Mentor mid-level and junior engineers; act as a technical escalation point

Partner with product, analytics, and platform teams to translate requirements into scalable systems

Own production reliability — on-call ownership, incident response, root-cause analysis for pipeline failures

Evaluate and introduce new tools/frameworks where they improve the system

Contribute to capacity planning and cost optimization for cluster infrastructure

Required Qualifications

Bachelor’s or Master’s degree in Computer Science, Engineering, or related field

7+ years of professional Java development experience

5+ years hands-on experience with Apache Spark in production environments

Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management

Proven track record designing systems processing terabyte+ scale data

Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg)

Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark

Proficiency with Kafka

Strong grasp of CI/CD, containerization, and infrastructure-as-code practices

Preferred Qualifications

Experience with Flink or other stream-processing frameworks

Familiarity with data governance, lineage, and quality frameworks

Experience with workflow orchestration at scale

Background in system design for multi-tenant or multi-region data platforms

Prior experience leading a team or acting as a technical lead

Show full description