Data Engineer
Коротко
Designs and builds on-premises data orchestration platform with Kubernetes and open-source tools. · Needs: Bachelor's degree in CS or related field; 5+ years data engineering, 2+ years enterprise-scale; workflow orchestration expertise.
Описание от работодателя
About the Role
We're seeking a talented Senior Data Engineer to join our team in a cross-cutting role that will help define and implement our next-generation data platform. In this pivotal position, you'll lead the design and implementation of scalable, self-service data pipelines with a strong emphasis on data quality and governance. This is an opportunity to shape our data engineering practice from the ground up, working directly with key stakeholders to build mission-critical ML and AI data workflows.
Key Responsibilities
Design, build, and maintain our on-premises data orchestration platform using the bestin-breed open source tools
Create self-service capabilities that empower teams across the organization to build and deploy data pipelines without extensive engineering support
Implement robust data quality testing frameworks that ensure data integrity throughout the entire data lifecycle
Establish data engineering best practices, including version control, CI/CD for data pipelines, and automated testing
Collaborate with ML/AI teams to build scalable feature engineering pipelines that support both batch and real-time data processing
Develop reusable patterns for common data integration scenarios that can be leveraged across the organization
Work closely with infrastructure teams to optimize our Kubernetes-based data platform for performance and reliability
Mentor junior engineers and advocate for engineering excellence in data practices
Qualifications
At least a Bachelor’s degree in Computer Science, Computer Engineering, Data Science, Mathematics, Engineering, or related disciplines.
5+ years of professional experience in data engineering, with at least 2 years working on enterprise-scale data platforms
Deep expertise with orchestrating workflows, performance optimization, and operational management
Strong understanding of data transformation techniques, including experience with testing frameworks and deployment strategies
Experience with stream processing frameworks and technologies
Proficiency with SQL and Python for data transformation and pipeline development
Familiarity with containerized application deployment
Experience implementing data quality frameworks and automated testing for data pipelines
Ability to work cross-functionally with data scientists, ML engineers, and business stakeholders
Preferred Qualifications
Experience with self-hosted data orchestration platforms (rather than managed services)
Background in implementing data contracts or schema governance
Knowledge of ML/AI data pipeline requirements and feature engineering
Experience with real-time data processing and streaming architectures
Familiarity with data modeling and warehouse design principles
Prior experience in a technical leadership role
We emphasize building systems that are maintainable, scalable, and focused on enabling selfservice data access while maintaining high standards for data quality and governance.
About You
The ideal candidate is a problem-solver who enjoys working on complex data systems and is passionate about data quality. You thrive in collaborative environments but can also work independently to deliver solutions. You're comfortable working directly with technical and nontechnical stakeholders and can communicate complex technical concepts clearly. Most importantly, you're excited about creating systems that empower others to work with data efficiently and confidently.