Senior Site Reliability Engineer

ZZPICTDirecte opdracht
ZZPICTDirecte opdracht
Plaats
Amsterdam, Noord-Holland
Reageren tot
1 januari 2027
Opdrachtgever
Booking.com

Kies je bemiddelaar

Je reageert via

Booking.com

onbekendTransparantie
geenScore
onbekendOpdrachten
onbekendFee
Wie betaalt de fee
niet vastgesteld
Tarief bij deze opdracht
fee onbekend
Beoordelingen
nog geen

Omschrijving

The role

Booking.com runs large-scale experiments as a vital part of its software development cycle. The Experimentation Platform enables product teams to make data-driven decisions by safely assigning experiments, collecting tracking data and providing statistically reliable metrics across user interactions.

We are looking for a Senior Site Reliability Engineer I to treat operations as a software problem and strengthen the reliability of the Experimentation Data Platform. You will work across services, streaming infrastructure and data pipelines, with a particular focus on a high-risk platform migration.

You will design and implement complex technical solutions, lead incident response for issues affecting the team, reduce human toil through automation and help establish reliable engineering practices. This is a hands-on contractor role: production experience with the relevant languages, streaming systems and cloud infrastructure is required so that you can contribute immediately.

What you will do

Build software and automation in Java and/or Python that improves availability, scalability, latency, efficiency and operational safety; write readable, reusable code and guide less-experienced engineers in these practices.

Design reliable solutions for distributed services and data pipelines, evaluating cost, business requirements, non-functional requirements, technology options and future extensibility.

Own systems end to end by monitoring application health and performance, setting relevant service and business metrics, supporting deployment and operations, and maintaining runbooks and operational documentation.

Lead incident response for team-owned issues, mitigate customer impact within agreed service levels, contribute to postmortems and deliver long-term fixes through root-cause analysis.

Reduce toil and operational cost by removing bottlenecks, addressing technical debt, preparing for scale and automating repetitive operational work.

Improve monitoring and alerting by reviewing observability metrics, business KPIs and capacity signals, and partnering with development teams to define useful reliability indicators.

Advise product and engineering teams on architecture, communicate clearly with stakeholders, challenge assumptions constructively and coach colleagues on reliability practices.

What you bring — required

Approximately 5–8 years of relevant engineering experience, or equivalent practical experience; a master’s degree or equivalent professional experience is welcome.

Strong professional programming experience in Java and/or Python, including debugging, testing, refactoring and safely changing production code. Experience in both languages is preferred.

Hands-on experience operating distributed production systems or data platforms, including Apache Kafka and Apache Flink.

Practical experience with Kubernetes and AWS, including EKS, S3 and IAM or equivalent cloud infrastructure.

Experience operating streaming pipelines and reliable file-based handoffs, including Parquet on object storage.

Strong experience with observability, alerting, incident response, root-cause analysis, postmortems and production documentation.

Experience designing or supporting high-risk production migrations, including rollout, rollback, reconciliation or backfill strategies.

Demonstrated ability to design complex technical solutions, identify underlying issues, improve engineering processes and communicate decisions clearly.

Experience coaching, guiding or enabling less-experienced engineers and partner teams.

Particularly valuable

Experience with the Flink Kubernetes Operator or large-scale streaming data pipelines.

Experience with data freshness, completeness and correctness monitoring in addition to service availability.

Familiarity with downstream Spark/Airflow or governed data platforms is useful but not required; the team owns the upstream pipeline and its S3/Parquet handoff rather than BDX itself.

… lees de volledige omschrijving bij Booking.com.

Bekijk meer opdrachten

Vergelijkbare opdrachten