About this role
This role oversees and manages the ingestion framework to ingest data from various data sources for data analytics and AI purposes. Detailed activities include:
- Analyzing the data requirements across various entities, develop, implement and maintain data ingestion pipelines from source to data lake pipelines.
- Automating data validation steps and report generation.
- Automating codes/ scripts that can be repeatedly used across similar analytics and reporting requirements.
- Managing and ensuring accuracy and timeliness and automation solution in end to end data extraction and integration with analytics and management reporting system.
- Designing the jobs schedule in scheduling tools and developing the configuration files for job schedulers.
- Performing production L2 batch support after production deployment.
- Perform code review functions for applications / programs developed by team members.
- To be part of initiatives that brings data into the data lake and delivers insights.
- Monitor and measure performance to assure ongoing data ingestion is meeting the SLA and optimization of the ingestion process to manage the performance and the SLAs.
- Work effectively with other stakeholders such as data engineering team, IT team, etc.
- Troubleshoot MapReduce/Spark Jobs and do performance tuning in production environments.
- Independently develop and sustain technical knowledge, certifications, and skills.
- Effectively handling day-to-day assignments given moderate directions and supervision.
- Bachelor or Master Degree in Computer Science, Engineering, or similar relevant field.
- Working experience in data ingestion or data engineering with Hadoop tech stack for 8+ years.
- Proficient with Scala and PySpark.
- Hands-on experience on Spark framework and other distributed data processing frameworks like Hadoop Map-Reduce, Hive etc. Proficient in ETL tools like Talend.
- Proficient in RDBMS databases such as Oracle, MySQL, MSSqlServer.
- Strong scripting skills in Linux environment and SQL.
- Expertise in Hadoop ecosystems.
- Hands-on Experience in Sqoop, Hive, Spark, Python, Scala is a must.
- Hands-on Experience in Job orchestration / Job schedulers like Autosys, Control-M
- Good to have working experience with one of the cloud platforms like AWS (Amazon Web Services), Microsoft Azure or Google Cloud Platform.
- Ability to plan and organize technical work and deliverables.
Ability to follow guidelines and adhere to the established software development standards and conventions.
S
elf-motivated and independent.
- Able to work with minimum supervision and to work well with stakeholders and project staff.
Ability to prioritize and multi-task across numerous work streams.
S
trong interpersonal skills; ability to work on cross-functional teams.
- Strong verbal and written communication skills.
- Deep knowledge of best practices through relevant experience across data-related disciplines and technologies particularly for enterprise wide data architectures and data warehousing/BI.
- Demonstrated problem-solving skills. Ability to learn effectively and meet deadline.
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?