As Lead Site Reliability Engineer, you’ll own the reliability strategy and architectural evolution of Vinted’s ML platform. Your mission is to ensure that every machine learning model—from classic algorithms to Large Language Models (LLMs)—is trained and served with industry-leading stability and performance.
You’ll join the ML Platform team, which owns the tooling for ML/LLM development, deployment and other platform capabilities that increase ML/AI delivery speed at Vinted. You won’t just be managing services; you’ll be shaping the architectural strategy for how Vinted serves AI at scale, and you'll act as a multiplier for the engineering function by mentoring senior engineers and driving domain-wide standards.
In this role, you’ll partner closely with Data Scientists and Platform teams to translate Vinted’s ML innovation into production reality. You will be responsible for orchestrating one of the region’s largest GPU fleets—over 200 high-performance GPUs dedicated to ML model training, production inference and LLM serving, including next-generation hardware designed for the most demanding AI workloads.
Our tech stack: Kubernetes, Terraform, Chef, Google Cloud Platform, Go, Kafka, Vespa, Redis, Vitess.
Nice to have
Nuoroda į skelbimą bus pridėta automatiškai žinutės pabaigoje.