Staff , Site Reliability Engineer - Cloud Platform

Butterfly Network, Inc. · Worldwide

Apply ↗full time

Company Description

Butterfly Network, Inc. (NYSE: BFLY) is driving a digital revolution in ultrasound imaging and sensing with its proprietary Ultrasound-on-Chip™ semiconductor technology and software solutions. Butterfly first proved its technology in the point-of-care ultrasound market – commercializing the world’s first single-probe, whole-body portable ultrasound device, which is now on its best-selling, third-generation: Butterfly iQ3™. The Company combines its advanced hardware with cloud software and AI, an enterprise workflow solution (Compass AI™) and other offerings to drive adoption of affordable, accessible ultrasound. Butterfly also enables third-party development of imaging AI apps through Butterfly Garden™, its software development kit and AI partnership initiative.

In addition to its medical imaging products, Butterfly Embedded™ is the Company’s Ultrasound-on-Chip™ licensing and co-development program designed to enable a new wave of ultrasound-enabled technologies across non-competitive healthcare markets and beyond. Through Butterfly Embedded™, partners can build and scale novel ultrasound applications powered by Butterfly’s proprietary semiconductor chip and software platform. Butterfly’s innovations have been recognized by Prix Galien USA, Fierce 50, TIME’s Best Inventions and Fast Company’s World Changing Ideas, among other achievements.

We’re a team of bold thinkers, problem-solvers, and innovators ready to shape the future of medical imaging. Let’s build something extraordinary together!

Job Description

We’re looking for a Staff Site Reliability Engineer to raise the bar for how we observe and operate our production cloud platform. Our cloud platform powers clinical workflows on top of one of the largest ultrasound image repositories in the world. In this role you’ll be a hands-on technical contributor who leads by execution.

What you’ll do

  • Own observability end to end. Evolve and streamline our metrics, logs, and distributed-tracing strategy. Building telemetry standards that give every team a clear, real-time picture of production health
  • Drive reliability through SLOs. Establish meaningful service-level objectives with product and engineering teams, and use error budgets to balance velocity against stability.
  • Lead incident response. Help set the standards for our incident response feedback loop. Manage the on-call rotations for production systems, respond to incidents with calm and rigor, manage blameless postmortems – then systematically reduce toil.
  • Engineer for reliability. Operate and improve our Kubernetes/EKS workloads on AWS. Build automation and tooling that removes manual operational work and improves developer productivity.
  • Be a force multiplier. Mentor engineers on operational excellence, set standards for how services are built to be observable and reliable, and influence architecture decisions across teams.

Qualifications

Baseline skills/experiences/attributes:

  • 8+ years of managing production systems
  • Deep, hands-on production experience with AWS and Kubernetes
  • Strong programming/scripting skills and experience building operational automation
  • Proven ownership of a full stack observability platform (NewRelic/Datadog/etc)
  • A track record of leading incident response and delivering improvements in reliability
  • Clear written and verbal communication skills
  • Strong desire for being a mentor to others

Ideally, you also have these skills/experiences/attributes (but it’s ok if you don’t!):

  • Background operating in a highly regulated or compliance-sensitive environment (e.g., HIPAA, SOC 2, FedRAMP, SaMD, etc).
  • Background in technical Healthcare Protocols (DICOM, HL7, FHIR) and systems integrations (PAC…

Don't just apply to this one.

Cold email the hiring manager and skip the pile. I'll write the email for you.

Book a call ↗

Related jobs