Senior Data Engineer

Location
Bucharest, Romania
Workplace
On-site

About this role

Role Overview

We are seeking an experienced Senior Data Engineer to design, build, and maintain scalable data pipelines and cloud-based data solutions. The role covers data ingestion, processing, storage, orchestration, and delivery to Analytics and BI teams within an AWS-based data ecosystem.

Responsibilities

  • Design and maintain batch and real-time data pipelines.
  • Integrate data from APIs, databases, files, and external sources using Apache Kafka and AWS services.
  • Develop and orchestrate workflows with Apache Airflow, including monitoring, retry mechanisms, and alerting.
  • Build and optimize distributed data processing solutions using Apache Spark, PySpark, Spark SQL, and AWS EMR.
  • Manage Data Lake and Data Warehouse environments using Amazon S3 and Amazon Redshift.
  • Design analytical data models, including Star and Snowflake schemas.
  • Prepare and optimize datasets for Analytics and Business Intelligence use cases, particularly for Qlik.
  • Develop high-quality solutions using Python and SQL, following best practices for Git, code reviews, testing, and documentation.
  • Ensure data quality, security, governance, and compliance, including AWS IAM, encryption, and access control.
  • Monitor production data pipelines and proactively address operational and performance issues.

Mandatory Requirements

  • Proven experience as a Senior Data Engineer or in a similar role.
  • Advanced proficiency in Python and SQL.
  • Hands-on experience with:
    • Apache Spark / PySpark
    • Apache Kafka
    • Apache Airflow
  • Strong AWS experience, particularly with:
    • Amazon S3
    • Amazon Redshift
    • AWS EMR
  • Experience with Data Lakes, Data Warehouses, ETL/ELT processes, and data modeling.
  • Strong understanding of batch and streaming architectures.
  • Experience with performance optimization, Git, testing practices, and production support.

Technical Stack

  • Python
  • SQL
  • Apache Spark / PySpark
  • Apache Kafka
  • Apache Airflow
  • AWS EMR
  • Amazon S3
  • Amazon Redshift
  • Git
  • Qlik

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

Know someone who'd be great for this?