DevOps Engineer
Коротко
Owns observability for critical trading paths, builds CI/CD and self-service platforms, and leads technical projects for platform reliability and developer · Needs: 5+ years software engineering including SRE/DevOps; deep expertise in observability (Datadog preferred).
Описание от работодателя
Senior DevOps Engineer - Remote We're seeking a Senior DevOps Engineer to join our team remotely. This role is perfect for a seasoned professional who is passionate about establishing robust quality, reliability, and automation practices platforms. You will be instrumental in designing, building, testing, and supporting infrastructure, CI/CD, and self-service platforms that make releases and environment management predictable and reversible.
Key Responsibilities Own and improve observability for critical trading paths through centralized logging, metrics, tracing, alerting, and runbooks to shorten mean-time-to-detect and mean-time-to-recover.
Create documentation and resources to enable engineers to use platform tools effectively.
Work closely with Technical Leadership to plan and prioritize platform reliability and developer-experience initiatives.
Partner with Core and Ops on production readiness, including capacity, failover, environment parity, and safe operational workflows.
Develop a robust Continuous Delivery practice within the team, prioritizing rapid feedback loops without increasing risk to live trading.
Lead technical projects including architecture and design decisions, code reviews/pairing, and mentoring of less experienced engineers.
Assist with hands-on delivery and secondment-style support when rolling out platform changes across teams.
Requirements / Skills You should bring 5+ years of software engineering experience , including time as a Site Reliability Engineer (SRE) or DevOps, with a strong focus on observability and production reliability. Excellent communication skills—written and verbal—are crucial, as you'll need to present ideas clearly and influence technical and non-technical stakeholders.
Deep expertise in observability practices including centralized logging, distributed tracing, and metrics collection—preferably Datadog, or familiarity with Prometheus, Grafana, OpenTelemetry, ELK, and similar stacks.
Hands-on experience designing, implementing, and maintaining logging pipelines and monitoring strategies for reliable, scalable, secure systems.
Proficient with Infrastructure as Code and automation tools such as Terraform, CloudFormation, or CDK.
Demonstrated experience designing and managing cloud-native applications and services on AWS.
Strong understanding of DevOps practices, automation, orchestration, containers, and serverless architectures.
Experience supporting Java / JVM production services is a strong plus.
Familiarity with Continuous Delivery practices including CI, test automation, and safe deployment strategies.
Comfortable working with distributed, global teams and asynchronous collaboration.
Adept at making rapid, high-quality decisions in fast-paced environments, with a bias toward action and ownership.
If you're committed to continuous learning, especially in observability, reliability engineering, and developer experience, we would love to hear from you.