Data Engineer
indrive
Job description
About the role
We are seeking a Data Engineer to join our Data Platform team, supporting Marketing, Growth, Partner and Finance data domains. You will work with cutting‑edge cloud technologies (GCP, AWS, BigQuery, Databricks, Kubernetes) to build large‑scale data infrastructure for analytics, machine learning and streaming data delivery.
Key responsibilities
- Build and operate batch and streaming ingestion pipelines into a layered BigQuery data warehouse using Airflow, Debezium CDC over Kafka, protobuf, Pub/Sub and Dataflow.
- Integrate external data sources such as GA4, AppsFlyer, TikTok/Meta/Google Ads, payment providers, S3 buckets and third‑party APIs, handling schema contracts, backfills and reconciliation.
- Engineer the data platform in Python, creating custom Airflow operators, Kafka Connect on Strimzi, Cloud Functions and API connectors.
- Develop CI/CD and change‑management tooling for BigQuery, including GitHub‑based test‑and‑approval flows, SQL migration engines (Liquibase/Flyway/Bytebase), sandbox validation, backup and rollback.
- Ensure reliability and correctness of pipelines – idempotency, deduplication, late‑data handling, backfills, freshness monitoring and alerting, with integration and unit tests.
- Drive data governance and compliance: ITGC‑compliant change management, IAM least‑privilege access, PII policy tags, DLP, Unity Catalog, column‑level lineage (OpenMetadata/Dataplex) and disaster‑recovery planning.
- Build internal data tools and platform services for agentic workflows – Streamlit apps, Slack bots, LLM‑based agents and MCP servers.
- Support analysts and business teams with data requests, fostering data‑driven decision‑making.
- Contribute to system design and architecture together with the development team.
Required profile
- Strong practical experience writing clean, well‑structured and tested Python code for services and data pipelines.
- Solid software design skills (OOP, modularity, design patterns) and ability to build reusable platform tools.
- Hands‑on experience operating services in cloud environments (GCP, AWS or similar), including CI/CD, containerisation, monitoring and alerting.
- Familiarity with Kubernetes and Terraform for infrastructure management.
- Proficiency with data‑warehouse technologies, especially BigQuery, and strong SQL skills.
- Clear communication skills for interacting with non‑technical analysts and business stakeholders.
- Proactive ownership of services and ability to contribute ideas to the team.
Required skills
- Python
- Airflow
- Debezium
- Kafka (including Strimzi)
- Protobuf
- Google Pub/Sub
- Dataflow
- BigQuery
- Google Cloud Platform (GCP)
- Amazon Web Services (AWS)
- Databricks
- Kubernetes
- Terraform
- SQL
- Streamlit
- CI/CD (GitHub)
- Cloud Functions
- Liquibase / Flyway / Bytebase
- OpenMetadata / Dataplex
Questions fréquentes
Why are you reporting this job?
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 3 апта бұрын
Expires 1 ай ішінде
19 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
indrive