Source description
About the role
Platform Reliability & SRE foundations
Own reliability of core data services: Trino, Iceberg, S3 / Ceph, Kafka, Kafka Connect, Schema Registry
Define and enforce SLIs/SLOs, error budgets, and on-call runbooks — solid SRE foundations are non-negotiable
Build full-stack observability with Prometheus and Grafana: metrics, dashboards, alerting pipelines, and anomaly detection
Manage and harden PostgreSQL clusters via Patroni for high-availability control-plane services
Kafka ecosystem — Connect & Schema governance
Operate and scale Kafka Connect clusters: connector lifecycle, offset management, dead-letter queues, and task rebalancing
Maintain the Schema Registry as the single source of truth for Avro/Protobuf/JSON schemas — enforce compatibility rules and schema evolution policies
Monitor consumer lag, connector throughput, and broker health via Prometheus JMX exporters and Grafana dashboards
Ensure end-to-end data contract integrity between producers and Iceberg/S3 consumers
Kubernetes, Kube-in-Kube & Crossplane
Operate production Kubernetes clusters (GKE/EKS + on-prem) — capacity planning, upgrades, PodDisruptionBudgets, resource quotas
Architect and manage Kube-in-Kube topologies to provide strong tenant isolation for data platform workloads — each team gets a dedicated virtual cluster without the overhead of a full physical cluster
Automate infrastructure and resource provisioning with Crossplane: define composite resources (XRDs) so data teams can self-serve Kafka topics, Trino namespaces, and S3 buckets through Kubernetes-native APIs
Maintain GitOps pipelines for platform deployment and configuration drift detection
Lakehouse architecture & cloud migration
Migrate from public cloud data warehouse to VeepeeCloud Iceberg-based lakehouse — managing coexistence, schema evolution, and time-travel
Architect resilient ingestion, transformation, and serving layers around Trino + S3
Optimize Trino query performance: memory limits, spilling, cost-based optimizer tuning
Agentic & developer enablement
Build agentic self-service tooling so data teams can provision Trino/Iceberg resources and Kafka Connect pipelines autonomously via Crossplane — reducing toil and ops bottlenecks
Develop FinOps dashboards (compute, storage, query cost) with Grafana and Prometheus-based cost exporters
Write clear technical documentation, runbooks, and internal ADRs
Multi-DC resilience & DRP
Design and implement multi-datacenter strategies across FR1 / NL1 — active-active and active-passive topologies
Leverage Fast Erasure Coding on object storage (Ceph/S3) to maximize durability with minimal replication overhead
Ensure data replication consistency across sites for Iceberg table metadata, Trino catalogs, and Schema Registry subjects
Lead DRP exercises: failover playbooks, RTO/RPO validation, postmortems
More at Veepee
Related open roles
Application - Support & Development Engineer (L2) - Freelance (W/M/X)
France · Hybrid
Responsable Excellence IA Finance (Finance AI Excellence Lead) - CDI (H/F/X)
France · Hybrid
BPO - Supply X AI Specialist
FR · Hybrid
Supply BPO x AI Specialist
Paris · Hybrid
Technicien de Maintenance H/F CDI
France · Onsite
SRE
France · Hybrid
