Required Skills: PySpark, Hive, Python, SQL, Unix, Kafka, Java
Job Description
Must Have Technical/Functional Skills
Primary Skill: PySpark, Hive, Python, SQL
Secondary: Unix, Kafka, Java
Roles & Responsibilities
We are seeking a highly experienced Hadoop Spark Developer with 10+ years of expertise in Big Data technologies, including PySpark, Hadoop Ecosystem, Hive, and Python. The ideal candidate will be responsible for designing, developing, optimizing, and maintaining large-scale data processing solutions.
Experience with Microsoft Copilot for AI-assisted development and productivity enhancement is highly desirable. The developer should hold a bachelor's or master's degree.
The candidate should possess strong analytical skills, hands-on experience in distributed data processing, and the ability to work closely with business stakeholders, architects, and data engineering teams.
- Design, develop, and maintain scalable Big Data solutions using Hadoop and Spark.
- Build and optimize ETL/ELT pipelines using PySpark, Hive, and Python.
- Process and analyze large datasets in distributed environments.
- Develop high-performance Spark jobs and optimize existing workloads.
- Create and manage Hive tables, partitions, views, and complex queries.
- Implement data quality, data validation, and reconciliation frameworks.
- Perform code reviews and ensure adherence to coding standards and best practices.
- Utilize Microsoft Copilot to accelerate development, automate code generation, troubleshooting, documentation, and testing activities.
- Strong experience building both batch and real-time streaming applications with Kafka.
- Collaborate with Data Architects, Data Scientists, Business Analysts, and DevOps teams.
- Troubleshoot production issues and perform root cause analysis.
- Design data ingestion frameworks for structured, semi-structured, and unstructured data.
- Participate in Agile ceremonies including sprint planning, estimation, and retrospectives.
- Mentor junior developers and provide technical leadership.