Source description
About the role
Own and evolve the core platform infrastructure: API serving layer, state management, policy enforcement engine, and execution runtime for agentic workflows.
Design and operate high-throughput, low-latency distributed services that back our model API products — including request routing, load management, rate limiting, and multi-tenant isolation.
Build and maintain downstream data pipelines (ETL/ELT) for API logs, usage analytics, and billing — ensuring data correctness, freshness, and queryability at scale.
Develop production-grade internal SDKs and libraries with clean APIs, strong type safety, and clear contracts that product teams can build on confidently.
Architect context and memory systems for conversational workloads — low-latency retrieval, caching, and integration with vector stores and retrieval pipelines.
Instrument end-to-end observability: define SLIs/SLOs, build structured logging and tracing, and drive reliability improvements across the platform.
Collaborate closely with ML and product teams to integrate model serving, voice runtime, and tooling infrastructure under tight latency and quality constraints.
More at Boson AI
