Senior Backend Infrastructure Engineer
Описание от работодателя
Required skills: Security, Java, Backend, Machine Learning (ML), Microsoft Platform, load balancing, Deployment, Debugging, Kubernetes, Transition
Role Summary
The Fabric squad builds the infrastructure that lets backend services find, trust, and talk to one another. The team owns service discovery, traffic routing, and service authentication across the production fleet. These are foundational and critical systems with fleet-wide impact, so reliability and safe rollout matter
We’re looking for a Senior Backend Infrastructure Engineer to help evolve this platform. You’ll work on systems such as Nameless, our service discovery platform; the xDS control plane that updates gRPC and Envoy clients; and the libraries that shape routing and service authentication. We’re looking for someone who can contribute hands-on, make sound architectural decisions, and help the team roll out changes safely across our large fleet.
We prioritize engineers who communicate clearly and collaborate well within and across teams, including in a remote setting.
The Opportunity
The Platform Mission creates the technology that enables the company to scale globally, experiment rapidly, and build for a billion users. Fabric sits within Core Infrastructure and provides the service-to-service foundation that other engineering teams build on. This is an opportunity to work on infrastructure with broad reach and challenging technical problems.
The team is expanding its traffic management capabilities, including developing capacity-aware routing and support for workloads beyond the existing gRPC mesh, while advancing the transition to SPIFFE and mTLS-based service authentication. You’ll work closely with partner teams in infrastructure, security, and machine learning to turn these capabilities into dependable production systems.
What you’ll work on
You’ll help build and harden the control plane and client libraries that connect our backend services. The work spans real-time service discovery, load balancing, regional failover, and service identity. It combines new capabilities with careful migrations and operational improvements to systems that are already used across the fleet.
You’ll help define rollout paths, improve observability, and make it easier for service owners to adopt the platform safely.
Key responsibilities
Design and implement changes to the xDS control plane and gRPC integrations that deliver endpoint and routing updates.
Develop traffic management capabilities, including zone-aware routing, regional failover, and capacity-aware spillover.
Strengthen metrics, debugging tools, and operational practices for critical production services.
Plan and execute fleet-wide migrations with clear validation and rollback paths.
Work with partner teams and service owners to make new capabilities practical to adopt.
Lead design reviews, contribute architectural proposals, and support teammates.
What we are looking for
5+ years of experience in backend or infrastructure engineering.
Strong Java experience and a track record of building and operating production services or shared libraries.
Solid understanding of distributed systems, including service discovery, load balancing and failure handling.
Experience operating production services on GCP and Kubernetes (GKE), including deployment, observability, and incident response.
Experience with gRPC and Kubernetes; familiarity with xDS, service meshes, or Envoy is a plus.
Strong debugging and performance skills, especially in systems with many clients and changing fleet state.
Experience rolling out infrastructure changes safely across multiple teams or services.
Familiarity with authentication, mTLS, or workload identity is a plus.
Clear communication and the ability to work effectively with remote teammates and partner teams.
Start/end: 2026-10-19 to 2027-04-18
Workplace: Remote within Sweden.