About this role
Owning quality for a high-scale,
low-latency streaming Customer Data Platform deployed on infrastructure we run
ourselves (on-prem, Kubernetes), not managed cloud services. The platform
ingests from thirty or more source systems at hundreds of thousands of events
per second, resolves customer identity, computes customer attributes in real
time, and emits governed signals to external destinations. You will build and
own the automation that proves this works: API and end-to-end test frameworks,
streaming and data-pipeline validation, latency and load testing against
explicit SLOs, and governance testing that proves consent, data-protection, and
tenant-isolation rules actually hold at runtime. Testing here is
evidence-producing, not exploratory only — test outcomes gate whether an
artifact is allowed to go live.
Requirements
• 6+ years
in software quality engineering, with the majority of your work in test
automation rather than manual testing
• Demonstrated
ownership of test automation frameworks you built or substantially
re-architected — not only writing test cases against someone else's framework
(core requirement)
• Ability to
write high-quality code in Python, Java, TypeScript, or equivalent languages,
including framework structure: base classes, shared utilities, configuration
handling, reporting, and test data management
• Strong API
test automation experience (REST Assured, Karate, pytest, Postman/Newman,
Playwright API, or similar) — test organization, authentication handling,
response and schema validation, and contract testing between services
• End-to-end
UI automation experience (Playwright, Cypress, Selenium, or similar) —
framework architecture, Page Object Model or equivalent, handling dynamic
elements, and parallel execution
• Demonstrated
ability to diagnose and eliminate flaky tests — proper waits, test isolation,
deterministic setup and teardown, and root-cause investigation rather than
blanket retries
• Performance
and load testing experience (JMeter, k6, Locust, Gatling, or similar) —
designing load scenarios, identifying bottlenecks, and validating throughput
and latency against stated targets
• Strong SQL
for data validation — joins, aggregations, window functions, and
source-versus-target reconciliation queries over large datasets
• Practical
experience testing data pipelines and ETL/streaming jobs — schema validation,
source-to-destination comparison at row and value level, completeness and
duplication checks, and sampling strategies for datasets too large to compare
in full
• Experience
testing event-driven systems and message queues (Kafka or equivalent) — event
and payload validation, ordering guarantees, delivery semantics (at-least-once
versus exactly-once-effective), idempotency, consumer lag, and offset behaviour
• Experience
testing stateful stream processing (Apache Flink or similar) — correctness
after restart from checkpoint, savepoint-based upgrades, state restoration,
late and out-of-order event handling, and behaviour under backpressure
• Ability to
validate end-to-end latency against an explicit SLO — measuring p95 and p99
across pipeline stages, and distinguishing in-scope platform processing from
excluded external calls such as third-party API round-trips and destination
acknowledgement
• Experience
integrating tests into CI/CD pipelines — trigger configuration, test stages,
artifact and report publication, and failure gates that block promotion
• Disciplined
test data management — setup and cleanup strategy, inter-test dependency
handling, and environment-specific data; experience working with synthetic,
generated, or masked datasets where production data cannot be used
• Ability to
produce structured, auditable test evidence: test results tied to the exact
artifact version and configuration under test, so that results are traceable
and cannot be silently reused after the artifact changes
• Experience
testing data governance and privacy controls — consent enforcement, data
classification and usage rules, tokenization and masking, and verifying that
raw sensitive identifiers do not appear in downstream topics, logs, or exports
• Experience
testing identity resolution or entity matching — deterministic and
probabilistic matching outcomes, merge and split behaviour, lifecycle state
transitions, and identifier reassignment scenarios
• Experience
testing multi-tenant systems — verifying tenant and workspace isolation, role-
and attribute-based access control, and absence of cross-tenant data leakage
• Experience
validating replay, backfill, and reconciliation — comparing recomputed values
against originally emitted values and confirming that reruns do not re-trigger
external side effects
• Experience
validating analytical or lakehouse data (Apache Paimon, Iceberg, Delta Lake, or
similar, queried through Trino, Spark, or equivalent) — comparing emitted
streams against materialized tables for completeness and correctness
• Experience
with failure-injection and resilience testing — node, broker, cache, or service
loss; verifying recovery path, data integrity after recovery, and absence of
duplicate or lost records
• Experience
testing on Kubernetes-deployed platforms — namespace-scoped environments,
Helm-based deployments, pod and job lifecycle, and access to logs and metrics
for diagnosis
• Working
familiarity with observability tooling for test diagnosis (Grafana, Prometheus,
OpenSearch or equivalent log search, distributed tracing)
• Experience
validating AI/ML or LLM outputs is an advantage — approaches to testing
non-deterministic responses, accuracy and regression measurement, and detection
of unsupported or fabricated output
• Rigorous
defect discipline — reproducible defect reports with evidence, clear severity
and triage judgement, and ownership of a maintained regression suite
• Domain
exposure to any of the following is an advantage: customer data platforms or
customer 360 systems, telecom or other high-volume transactional systems,
real-time systems with latency SLAs, AdTech or MarTech platforms, and data
analytics or dashboard systems
• Familiarity
using AI tools for development and debugging (Claude, Cursor, Codex)
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?