Principal / Staff Data Platform Engineer
Описание от работодателя
Required skills: SQL, Python, Artificial Intelligence (AI), CI/CD, Software as a Service (SaaS), Cloud, Scala, Spark, Amazon Web Services (AWS), Data architecture
Start: ASAP / Flexible
Contract length: 6 months + prolongations
Contract type: B2B
Remote: 100% remote
Business trips: none
The Principal / Staff Data Platform Engineer will lead the technical development of a new AI-first data foundation, focused on transforming governed events into trusted data products and analytics.
Main Responsibilities:
Lead the technical design of the new data platform.
Define data movement from the Iceberg-based event layer to analytical and operational data products.
Evaluate and select technologies for data querying, transformation, orchestration, storage, and serving.
Design canonical entities and reusable data models across various business domains.
Establish scalable patterns for batch and near-real-time data processing.
Make architectural decisions and establish engineering standards for the platform.
Build data quality, observability, lineage, and reconciliation into the platform.
Define engineering standards for testing, deployments, versioning, and schema evolution.
Mentor other data engineers as the team grows.
Key Requirements:
Significant experience in designing and operating production data platforms.
Strong experience with distributed data systems and high-volume event data.
Deep understanding of lakehouse architectures and open table formats like Apache Iceberg.
Strong SQL knowledge and experience in Python, Java, or Scala.
Experience with distributed processing technologies like Spark or Flink.
Strong understanding of data modelling, performance, and storage design.
Experience in building both batch and near-real-time data pipelines.
Experience with cloud infrastructure, preferably AWS.
Strong understanding of CI/CD and infrastructure-as-code principles.
Experience with data contracts and schema evolution.
Nice to Have:
Experience in ad-tech or high-volume event processing.
Multi-tenant SaaS data architecture knowledge.
Experience with semantic layers.
Experience building data platforms for analytics and ML/AI workloads.
Knowledge in privacy, residency, and regulated data environments.