About this role
This is a remote position.
Role Overview
We are looking for a Senior / Principal Data Architect to take technical ownership of a privacy-preserving data platform that enables organizations to securely transform and prepare operational data for AI/ML training and analytics without compromising individual privacy.
Key Responsibilities
- Design and lead the architecture of a scalable, secure, and cost-efficient AWS data lake/lakehouse platform using services such as Amazon S3, AWS Glue, Athena, Lake Formation, and related AWS data services.
- Define data architecture patterns for batch and real-time/CDC ingestion from transactional and operational systems.
- Design efficient Parquet-based data layouts, including partitioning strategies, file sizing, compaction, compression, and query optimization.
- Architect and implement data de-identification, tokenization, anonymization, and privacy-preserving techniques for sensitive datasets.
- Establish data governance frameworks covering data ownership, access control, classification, lineage, cataloging, retention, and auditability.
- Design robust data catalog and metadata management solutions while supporting schema evolution and data discovery.
- Define and enforce data quality, security, privacy, and compliance standards across the platform.
- Collaborate with Data Engineering, AI/ML, Security, Privacy, Legal, and Compliance teams to translate business and regulatory requirements into technical solutions.
- Optimize data platforms for performance, scalability, freshness, reliability, and cloud cost efficiency.
- Establish architecture and engineering best practices for teams building and consuming data products.
- Provide technical leadership and mentorship to data engineers and other technical stakeholders.
- Evaluate new technologies and architectural approaches to improve the platform's capabilities and support evolving AI/ML requirements.
Requirements
Required Skills & Experience
- 8+ years of experience in Data Engineering, Data Architecture, or related roles.
- 3+ years of experience designing and owning data platforms used by multiple engineering or analytics teams.
- Strong hands-on experience designing and implementing AWS data lake/lakehouse architectures.
- Strong knowledge of Amazon S3, AWS Glue, Amazon Athena, and AWS Lake Formation.
- Strong understanding of data modeling, data warehousing, data lakes, and lakehouse architectures.
- Hands-on experience with Parquet, partitioning, file optimization, compaction, and large-scale data processing.
- Experience designing CDC and batch data ingestion pipelines from transactional systems.
- Strong understanding of data cataloging, metadata management, data lineage, and schema evolution.
- Practical experience with data privacy, anonymization, de-identification, tokenization, or other privacy-preserving techniques.
- Strong understanding of data governance, access control, security, compliance, and regulatory requirements.
- Experience working directly with Privacy, Legal, Security, or Compliance teams.
- Strong understanding of data quality, data lifecycle management, and data retention.
- Experience with cloud cost optimization and performance engineering for large-scale data platforms.
Nice to Have
- Experience with GDPR, NDPA/DPA, or similar data privacy regulations.
- Knowledge of k-anonymity, differential privacy, or privacy-preserving analytics.
- Experience with synthetic data generation for AI/ML or analytics use cases.
- Experience with Apache Iceberg or other open table formats.
- Experience supporting AI/ML training data pipelines and feature engineering platforms.
- Experience with AWS security and identity services, including IAM and KMS.
- Experience with large-scale data cost optimization and FinOps.
- Experience designing multi-account or multi-environment AWS data platforms.
- Familiarity with infrastructure automation using Terraform or CloudFormation.
- Experience establishing data platform standards, reference architectures, and technical governance.
Required Skills
Data Architecture | Data Engineering | AWS | Amazon S3 | AWS Glue | Amazon Athena | AWS Lake Formation | Data Lake / Lakehouse | Data Modeling | Parquet | CDC | Data Ingestion | Data Catalog | Data Lineage | Schema Evolution | Data Governance | Data Privacy | Data De-identification | Data Anonymization | Data Tokenization | Data Security | Data Quality | AI/ML Data Platforms | Performance Optimization | Cloud Cost Optimization
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?