. Design and develop robust ETL/data pipelines using Hadoop ecosystem tools.
2. Build and optimize large-scale data processing jobs using Spark and/or Hive.
3. Develop workflows using Airflow.
4. Write complex SQL/HiveQL for data transformation and reporting needs.
5. Ingest structured and unstructured data from multiple sources (Kafka, Sqoop, APIs, files).
6. Ensure data quality, lineage, and governance standards are followed.
7. Monitor, troubleshoot, and tune jobs for performance and scalability.
8. Create technical documentation and follow coding best practices.