Key Responsibilities Design, develop and maintain scalable ETL/ELT data pipelines for large-volume structured and unstructured datasets. Develop high-performance data processing solutions using Python, PySpark, Apache Spark, Spark SQL and Scala . Build batch and real-time data pipelines using Kafka, Kinesis, Spark Streaming, AWS Glue and Airflow . Develop data ingestion frameworks integrating RDBMS, APIs, files, cloud storage and streaming platforms . Design and implement data lakes, data wareh…