Presto

Location
Bangalore North
Workplace
On-site

About this role

Job Description:

Responsibilities

  • ELT Pipeline Engineering: Design, build, and maintain high-throughput batch and streaming ETL/ELT data pipelines using SQL, Python, and PySpark to ingest large datasets into cloud data lakes and Snowflake
  • Lakehouse & Federated Query Operations: Manage and optimize data structures across AWS/Azure storage (S3/ADLS) using open table formats (Apache Iceberg, Delta Lake) and execute cross-platform queries via Presto/Trino
  • Query & Warehouse Optimization: Diagnose performance bottlenecks, optimize SQL queries, and configure Presto/Trino execution parameters alongside Snowflake virtual warehouse sizing/clustering to reduce compute costs and query latency
  • BI & Semantic Layer Development: Build and maintain scalable data models, Looker semantic layers (LookML), and curated Tableau data sources to enable self-service business intelligence and executive reporting
  • Governance & Access Control: Enforce data security and row/column-level access controls using AWS Lake Formation tag-based policies and Snowflake RBAC, while leveraging OpenSearch for real-time log monitoring and search
  • Data Quality & Orchestration: Implement automated pipeline testing, schema validation, and alert monitoring using dbt and Apache Airflow to maintain data freshness and platform integrity
  • Agile Delivery & Collaboration: Active participation in Agile sprint ceremonies, performing code reviews, creating technical architecture documentation, and collaborating with senior engineers on platform features

Qualifications

Technical Expertise:

  • Core Skills & Languages: Minimum 6 years of hands-on Data Engineering experience with expert-level proficiency in SQL and Python (Pandas, PySpark)
  • Engine & Warehouse Expertise: Strong operational knowledge of Snowflake and PrestoDB for complex, large-scale query processing
  • Cloud & Lakehouse Architecture: Practical experience with AWS or Azure services, cloud storage (S3/ADLS), and columnar/open table formats (Apache Iceberg, Delta Lake, Parquet)
  • BI & Analytics: Hands-on experience developing Looker models (LookML, PDTs) and constructing Tableau dashboards & reporting
  • Transformation & Streaming: Solid background in dbt for transformation, Apache Airflow for workflow DAG management, and familiarity with Apache Kafka for event streaming platforms
  • Applied Data Science Familiarity: Understanding of data preparation requirements for statistical models, machine learning frameworks, and regression/tree-based algorithms

Education & Experience:

  • Degree in Data Science, Analytics, Computer Science, or a related quantitative field
  • 6 years of progressive experience building, maintaining, and supporting production data platforms


Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

Know someone who'd be great for this?