Padmi
Veepee logo
Veepee

flash sales · online private sales

SRE - DataPlatform

France · HybridPosted 3 months ago
InfrastructureUnspecifiedPermanent
Apply at Veepee

Opens the source posting on jobs.lever.co

Source description

About the role

View original

Platform Reliability & SRE foundations

Own reliability of core data services: Trino, Iceberg, S3 / Ceph, Kafka, Kafka Connect, Schema Registry

Define and enforce SLIs/SLOs, error budgets, and on-call runbooks — solid SRE foundations are non-negotiable

Build full-stack observability with Prometheus and Grafana: metrics, dashboards, alerting pipelines, and anomaly detection

Manage and harden PostgreSQL clusters via Patroni for high-availability control-plane services

Kafka ecosystem — Connect & Schema governance

Operate and scale Kafka Connect clusters: connector lifecycle, offset management, dead-letter queues, and task rebalancing

Maintain the Schema Registry as the single source of truth for Avro/Protobuf/JSON schemas — enforce compatibility rules and schema evolution policies

Monitor consumer lag, connector throughput, and broker health via Prometheus JMX exporters and Grafana dashboards

Ensure end-to-end data contract integrity between producers and Iceberg/S3 consumers

Kubernetes, Kube-in-Kube & Crossplane

Operate production Kubernetes clusters (GKE/EKS + on-prem) — capacity planning, upgrades, PodDisruptionBudgets, resource quotas

Architect and manage Kube-in-Kube topologies to provide strong tenant isolation for data platform workloads — each team gets a dedicated virtual cluster without the overhead of a full physical cluster

Automate infrastructure and resource provisioning with Crossplane: define composite resources (XRDs) so data teams can self-serve Kafka topics, Trino namespaces, and S3 buckets through Kubernetes-native APIs

Maintain GitOps pipelines for platform deployment and configuration drift detection

Lakehouse architecture & cloud migration

Migrate from public cloud data warehouse to VeepeeCloud Iceberg-based lakehouse — managing coexistence, schema evolution, and time-travel

Architect resilient ingestion, transformation, and serving layers around Trino + S3

Optimize Trino query performance: memory limits, spilling, cost-based optimizer tuning

Agentic & developer enablement

Build agentic self-service tooling so data teams can provision Trino/Iceberg resources and Kafka Connect pipelines autonomously via Crossplane — reducing toil and ops bottlenecks

Develop FinOps dashboards (compute, storage, query cost) with Grafana and Prometheus-based cost exporters

Write clear technical documentation, runbooks, and internal ADRs

Multi-DC resilience & DRP

Design and implement multi-datacenter strategies across FR1 / NL1 — active-active and active-passive topologies

Leverage Fast Erasure Coding on object storage (Ceph/S3) to maximize durability with minimal replication overhead

Ensure data replication consistency across sites for Iceberg table metadata, Trino catalogs, and Schema Registry subjects

Lead DRP exercises: failover playbooks, RTO/RPO validation, postmortems

More at Veepee

Related open roles

View all roles
SRE - DataPlatform at Veepee · Padmi