About this role
- Design, develop, and maintain scalable ETL/ELT pipelines using Python, PySpark, and SQL.
- Perform data extraction, transformation, and loading from multiple source systems into enterprise data platforms.
- Develop reusable data transformation logic to support business reporting, analytics, and machine learning initiatives.
- Optimize SQL queries and PySpark jobs to improve performance and processing efficiency.
- Build and maintain data models, staging layers, and curated datasets for downstream consumption.
- Perform data cleansing, validation, reconciliation, and quality checks to ensure data accuracy and consistency.
- Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data-related issues.
- Collaborate with business analysts, data architects, and data scientists to understand data requirements and implement scalable solutions.
- Participate in code reviews, testing, deployment, and production support activities.
- Develop and maintain technical documentation, including data mappings, transformation logic, and ETL workflows.
- Ensure compliance with data governance, security, and regulatory standards.
- Monitor scheduled ETL jobs and proactively address operational issues.
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?