About this role
Owning the end-to-end solution architecture
of the Daitics AI CDP: a sovereign, on-prem, telco-native Customer Data
Platform built on a streaming architecture, deployed on infrastructure we run
ourselves (Kubernetes, Helm), not managed cloud services. The platform
processes hundreds of thousands of events per second across thirty or more
source systems and delivers unified customer profiles for B2C persons and B2B
organizations, accounts, sites, lines, devices and contacts. You will own the
architecture across the five planes (Authoring, Control, Data, Activation,
Observability) and the five layers (atomic events, tile primitives, trait
values, signal filters, signal events), hold the line on architectural
principles across engine boundaries, and translate them into designs
engineering teams can build against.
Requirements
• 12+ years
in software and data engineering, with at least 5 years owning solution or
platform architecture for large-scale distributed systems
• Demonstrated
ownership of end-to-end architecture for a streaming data platform on
self-managed infrastructure (on-prem or private cloud) topology, state
placement, failure domains, cutover, and capacity, not documentation alone
• Deep,
hands-on architectural command of Apache Flink as a stateful streaming runtime DataStream and SQL, keyed state on RocksDB, broadcast state for configuration
distribution, Async I/O for external lookups, checkpoint and savepoint
discipline, and failure isolation across independent job clusters (core
requirement)
• Proven
ability to design two-tier latency architectures: a millisecond-level hot path
against a minute-level durable cold path, with an explicit end-to-end latency
budget and a documented inclusion and exclusion list
• Experience
designing Kafka topologies for multi-tenant platforms topic taxonomy,
partition-key selection and partition sizing, compacted control streams,
retention classes, and enforced naming conventions
• Experience
architecting schema governance across producers and consumers Avro or
equivalent through a schema registry, compatibility modes, and coordinated
promotion of breaking changes
• Hands-on
architecture experience with lakehouse table formats (Apache Paimon, Iceberg,
Delta Lake, or Hudi) on S3-compatible object storage — primary-key and LSM
table design, partitioning, compaction, snapshot expiry, and separation of
streaming from batch read paths
• Experience
designing governed single-writer patterns where many producers write to one
substrate through one consolidating service, with per-stream configuration
applied without restarts and exactly-once-effective semantics
• Experience
with distributed in-memory caches (Apache Ignite, Hazelcast, Redis, or similar) partitioned versus replicated tables, affinity colocation with the compute
layer, thin-client access patterns, operator-local caching, and cross-consumer
invalidation
• Ability to
define and enforce a state-placement decision framework: broadcast state versus
keyed state versus external cache versus first-seen load, with explicit rules
rather than case-by-case judgment
• Experience
designing aggregation architectures that serve arbitrary rolling windows with
sub-millisecond reads while keeping state bounded and recomputable
(pre-aggregated tiles, multi-resolution granularity ladders, or equivalent),
including approximate aggregation through sketches (HyperLogLog, t-digest,
count-min) and honesty about their limits
• Experience
architecting identity resolution at scale deterministic and probabilistic
matching, identity graph modeling, effective-dated B2B relationships, merge and
split workflows, anonymous-to-known stitching, and downstream propagation of
identity change
• Experience
modeling entity lifecycle state machines in streaming systems, including
identifier reassignment, quarantine, archival, and purge, with auditable
transitions
• Architecture
experience with consent enforcement and data governance configurable
enforcement modes, data usage labeling and enforcement, policy distribution to
the runtime, and enforcement designs that avoid a synchronous service call per
event
• Experience
architecting tokenization, hashing, masking, and encryption for sensitive
identifiers — reversible versus irreversible choices, key versioning and
rotation (HashiCorp Vault Transit or equivalent), protection applied at the
ingest boundary, and audited just-in-time detokenization at delivery
• Experience
designing end-to-end lineage and edit-time impact analysis artifacts and
dependencies modeled as a graph, forward and reverse traversal, and prevention
of orphaned dependencies on delete, rename, or semantic change
• Experience
designing versioned artifact activation and cutover across engine boundaries
grace windows, transition emission policy, late-event windows, and rollback
• Ability to
reason about state reconciliation when a computation definition changes forward-only, full or partial recalculation, parallel versions, evacuate and
rebuild, snapshot seeding and how data retention constrains which options are
valid
• Experience
designing multi-tenant isolation on Kubernetes namespace and resource
topology per tenant and workspace, noisy-neighbour control, per-tenant storage,
secrets, and cross-workspace sharing contracts
• Experience
architecting authentication, authorization, and network security enterprise
IdP federation (OIDC, Keycloak or equivalent), attribute-based access control,
Kubernetes NetworkPolicies for east-west isolation, service-mesh mTLS, and
secrets management
• Experience
defining failure domains, SLOs, RPO and RTO targets for streaming platforms,
and designing replay, backfill, and reconciliation paths that do not re-trigger
external actions
• Ability to
size a platform of this shape end to end: from events per second and subject
counts to partitions, parallelism, keyed state size, cache footprint, and
storage — and to validate the model against measured behaviour
• Judgement
on when to wrap an open-source component behind a stable platform service
interface with pluggable implementations, and when to depend on it directly
• Experience
with a canonical or industry semantic model and the discipline of preventing
semantic drift; telco experience with TM Forum SID and TMF620 is a strong
advantage
• Experience
integrating ML into a streaming platform models upstream of derived
attributes, online and batch scoring paths, training-serving consistency, and
keeping inference off the hot path unless benchmarked
• Practical
experience with Spring Boot on Java for control-plane microservices, PostgreSQL
as a definition store, and workflow orchestration (Temporal, Argo Workflows, or
equivalent) with BPMN approval flows (Flowable, Camunda, or equivalent)
• Experience
with observability stacks (OpenTelemetry, Prometheus, Grafana, distributed
tracing) and with defining what operators actually need to see
• Familiarity
with graph and vector substrates for lineage, relationship traversal, and
semantic similarity over platform metadata
• Ability to
author and maintain a layered architecture corpus a master blueprint with
component annexes, explicit design rules, and cross-document consistency and
to defend it in architecture review
• Familiarity
using AI tools for design, development, and debugging (Claude, Cursor, Codex)
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?