Padmi
Parallel Wireless logo
Parallel Wireless

Open RAN · 5G/4G/3G/2G software

Senior/Principal Local LLM & Generative AI Platform Engineer

Israel · HybridPosted 15 days ago
Machine learningStaff+Full Time
Apply at Parallel Wireless

Opens the source posting on jobs.lever.co

Source description

About the role

View original

Own the architecture and technical roadmap for a secure, reliable, and maintainable local LLM platform deployed in Parallel Wireless-controlled infrastructure.

Partner with engineering, product, support, IT, information security, legal, and domain experts to prioritize high-value use cases and translate them into measurable product and platform requirements.

Build a modular inference and model-gateway layer with stable APIs, model routing, streaming, concurrency controls, quotas, and the ability to change models or serving backends without rewriting every application.

Evaluate open-weight language, code, embedding, reranking, and, where useful, multimodal models against PW-specific tasks; document model provenance, licenses, limitations, security posture, hardware needs, and total cost of ownership.

Optimize serving across available CPU, GPU, and accelerator resources using techniques such as continuous batching, caching, parallelism, quantization, and right-sized context limits while protecting output quality.

Design and operate RAG and enterprise-search pipelines for approved repositories, wikis, tickets, standards, design documents, test results, logs, and support content, including parsing, chunking, metadata, embeddings, hybrid retrieval, reranking, freshness, citations, and deletion.

Enforce source-system permissions throughout ingestion and retrieval so that the platform never exposes content a user is not authorized to access; integrate with company identity, SSO, role-based access control, secrets management, and audit logging.

Establish versioned evaluation datasets and automated offline and online evaluation for retrieval quality, groundedness, factual accuracy, citation quality, code correctness, task completion, latency, safety, and refusal behavior.

Create release gates and reproducible regression tests for changes to models, prompts, tools, embeddings, retrieval logic, indexes, and serving configurations; support canary releases, rollback, and clear approval paths.

Implement end-to-end observability for model and agent workflows, including traces, errors, time to first token, inter-token latency, throughput, queue time, resource utilization, saturation, availability, and user feedback.

Design safe tool-calling and agent workflows with least-privilege access, sandboxing, input and output validation, bounded execution, human approval for consequential actions, and complete traceability.

Integrate the platform into the tools employees already use—such as developer environments, source-control and CI workflows, knowledge systems, ticketing systems, and internal applications—through reusable SDKs, APIs, and reference implementations.

Build the operational foundations for production use: CI/CD, configuration and model registries, backups, disaster recovery, capacity planning, dependency and vulnerability management, incident response, and lifecycle policies for models and data.

Protect proprietary and personal information through network isolation, encryption, retention controls, redaction where appropriate, secure logging, and defenses against prompt injection, data poisoning, unsafe output handling, and model-supply-chain risks.

Determine when prompt or retrieval improvements are sufficient and when parameter-efficient fine-tuning, distillation, or other adaptation is justified by measured quality gains.

Make the platform usable beyond the core AI team through documentation, examples, training, office hours, and hands-on collaboration; use telemetry and structured feedback to improve adoption and effectiveness.

Communicate architecture decisions, quality evidence, risk, capacity, and roadmap tradeoffs clearly to technical and business stakeholders.

More at Parallel Wireless

Related open roles

View all roles
Senior/Principal Local LLM & Generative AI Platform Engineer at Parallel Wireless · Padmi