About this role
Building high-scale, low-latency streaming
data pipelines deployed on infrastructure we run ourselves (on-prem), not
managed cloud services. You will design and operate high-volume real-time data
systems end to end, with Apache Flink as the core stream-processing engine.
Requirements
• 7+ years
of experience in data engineering and software development
• Ability to
write high-quality code in Java/Scala, Python, or equivalent languages
• Deep,
hands-on production experience with Apache Flink — DataStream API and Table API
/ Flink SQL (core requirement)
• Demonstrated
experience with Flink state management: keyed state, state backends (e.g.,
RocksDB), large state sizes, and state TTL
• Hands-on
experience with checkpointing, savepoints, and fault tolerance — exactly-once
vs. at-least-once semantics, recovery, and savepoint-based job upgrades
• Strong
grasp of event-time processing: watermarking, windowing strategies, allowed
lateness, and late-data handling
• Experience
diagnosing and resolving backpressure — parallelism, operator chaining, and
network buffer tuning
• Experience
operating Flink on self-managed infrastructure (Kubernetes or YARN) —
application vs. session mode, high availability, and rolling upgrades
• Practical
experience with stream processing (Kafka Streams or equivalent) and messaging
systems for high-volume workloads, including exactly-once sinks and schema
registry usage
• Practical
experience with distributed query engines (e.g., Trino/Presto or similar)
• Practical
experience with ETL / data integration tools, commercial or open-source (e.g.,
Datastage, Informatica, Apache NiFi, or similar)
• Practical
experience with SQL-based transformation frameworks (e.g., dbt or others)
• Strong SQL
skills and understanding of data modeling and data warehousing for analytical
workloads
• Hands-on
experience with real-time / low-latency analytical stores (columnar or OLAP
engines, e.g., Apache Pinot/ClickHouse or similar)
• Practical
experience with big-data platforms and distributions (e.g., Cloudera, Hadoop
ecosystem, or similar)
• Practical
experience containerizing and operating data workloads (Docker; Kubernetes a
plus)
• Experience
with workflow orchestration tools (e.g., Airflow or similar)
• Familiarity
with data lake table formats (e.g., Apache Iceberg or similar), including
streaming ingestion, compaction, and small-file management
• Familiarity
with data governance / cataloging tools (e.g., DataHub or similar)
• Familiarity
with lakehouse management systems (e.g., Apache Amoro or similar)
• Familiarity
using AI tools for development and debugging (Claude, Cursor, Codex)
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?