DevOps Engineer (Rust) – Logfile Management,...
Noord-Holland · KPN (Noord-Holland) · Bij Oranje
Omschrijving
KPN DATA SERVICES HUB / VACANCY / JULY 2026 DevOps / Site Reliability Engineer Logfile Management KPN’s Logfile Management (LFM) team builds and operates the pipelines that transport, process and store logs and metrics at scale on the KPN Data Services Hub. These services support external customers and internal KPN teams in areas such as security (SIEM), observability and compliance. The team has grown rapidly, welcoming 12 new customers in the past quarter. To support this growth, KPN is looking for a DevOps/Site Reliability Engineer who can take its operational maturity to the next level. This is not primarily a software development role. The main focus is operating critical services: reliability, monitoring, incident response, certificate management, compliance and continuous operational improvement. You should genuinely enjoy taking ownership of these responsibilities. The team You will join a small team within KPN IoT Data, currently consisting mainly of software engineers. Your role will complement the team by bringing a strong operational and reliability-focused perspective. The team operates everything it builds and is responsible for the full service lifecycle. The observability platform is an important part of the service and is also offered commercially to customers. What you’ll do Take ownership of the operational processes surrounding KPN’s log and metrics services. Improve the reliability, scalability and operational maturity of critical services. Monitor, troubleshoot and improve services using the Grafana LGTM stack: Loki, Grafana, Tempo and Mimir. Improve monitoring, alerting, incident response, automation and CI/CD processes. Manage TLS/mTLS certificates for proxies, including the scheduled replacement of certificates approximately every six months. Deploy, operate and continuously improve services throughout their full lifecycle. Contribute to KPN’s compliance with standards such as ISAE 3000 and SOC 2. Participate in the 24/7 on-call rotation. Work closely with the software engineers and help the team make well-considered operational improvements. Contribute more broadly to the department where your experience and interests add value. Must-haves A genuine operational mindset: you gain energy from running and improving critical services rather than primarily building software. Experience in a DevOps, Site Reliability Engineering or comparable operational engineering role. Hands-on experience with the Grafana LGTM stack—Loki, Grafana, Tempo and Mimir—or comparable observability and monitoring tooling. Experience with CI/CD and DevOps ways of working. Experience with certificate management, including TLS/mTLS certificates for proxies. Good knowledge of networking fundamentals, including DNS, routing and firewalls. A proactive and ownership-driven mindset. An entrepreneurial attitude and the ability to identify and initiate improvements. Strong collaboration skills: you can present improvement proposals, discuss them constructively with engineers and bring the team along without forcing your views. Professional fluency in English, as English is the team’s working language. Willingness to participate in a 24/7 on-call rotation. Based in the Netherlands and able to attend the required office days. Availability for a long-term engagement, preferably for at least one year. Nice to have Experience operating workloads on Kubernetes. Experience with Git, GitHub Actions and Harbor. Software engineering skills, ideally in Rust or Scala.
… lees de volledige omschrijving bij Bij Oranje.
Je wordt doorgestuurd naar de website van Bij Oranje. ZZPdock is geen tussenpartij.