Fluidstack Logo

Fluidstack

Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
In-Office or Remote
33 Locations
Mid level
In-Office or Remote
33 Locations
Mid level
Site Reliability Engineers at Fluidstack ensure the reliable operation of global GPU clouds through hands-on infrastructure management and automation, deploying and optimizing solutions for AI workloads.
The summary above was generated by AI
About Fluidstack

Fluidstack is building GPU supercomputers for top AI labs, governments, and enterprises. Our customers include Mistral, Poolside, Black Forest Labs, Meta, and more.

Our team is small, highly motivated, and focused on providing a world class supercomputing experience. We put out customers first in everything we do, working hard to not just win the sale, but to win repeated business and customer referrals.

We hold ourselves and each other to high standards. We expect you to care deeply about the work you do, the products you build, and the experience our customers have in every interaction with us.

You must work hard, take ownership from inception to delivery, and approach every problem with an open mind and a positive attitude. We value effectiveness, competence, and a growth mindset.

About the Role

SREs at Fluidstack sit at the core of our infrastructure, working across software, hardware, and operations to ensure the reliability and performance of our global GPU cloud.

They partner closely with teams including networking, platform engineering, and data center operations to build systems that scale with the demands of AI workloads.

SREs are hands-on and possess deep systems knowledge and strong communication skills. You’ll be responsible for tackling complex production issues, deploying resilient infrastructure, and continuously improving the stability and observability of our platform as we grow.

A typical day may involve:

  • Deploying clusters of 1,000+ GPUs using custom written playbooks; modifying these tools as necessary to provide the perfect solution for a customer.

  • Validating correctness and performance of underlying compute, storage, and networking infrastructure, and working with providers to optimize these subsystems.

  • Migrating petabytes of data from public cloud platforms to local storage, as quickly and cost effectively as possible.

  • Debugging issues anywhere in the stack, from “this server’s fan is blocked by a plastic bag” to “optimizing S3 dataloaders from buckets in different regions”.

  • Building internal tooling to decrease deployment time and increase cluster reliability, including automation where the customer benefits clearly outweigh the implementation overhead.

This role will involve being part of an on-call rotation up to one week per month.

Focus
  • A customer-centric attitude, an accountability mindset, and a bias to action.

  • A track record of shipping clean, well-documented code in complex environments.

  • An ability to create structure from chaos, navigate ambiguity, and adapt to the dynamic nature of the AI ecosystem.

  • Strong technical and interpersonal communication skills, a low ego, and a positive mental attitude.

An ideal candidate meets at least the following requirements:

  • 2+ years of SRE, DevOps, Sysadmin, and/or HPC engineering experience.

  • Great verbal and written communication skills in English.

  • Experience deploying and operating Kubernetes and/or SLURM clusters.

  • Experience in writing Go, Python, Bash.

  • Experience using Ansible, Terraform, and other automation or IAC tools.

  • Strong engineering background, preferably in Computer Science, Software Engineering, Math, Computer Engineering, or similar fields.

Exceptional candidates have one or more of the following experiences:

  • You have built and operated an AI workload at 1000+ GPU scale.

  • You have built multi-tenant, hyperscale Kubernetes based services.

  • You have physically deployed infrastructure in a datacenter, managed bare metal hardware via MaaS or Netbox, etc.

  • You have deployed and managed multi-tenant InfiniBand or RoCE networks.

  • You have deployed and managed petabyte scale all-flash storage systems, including DDN, VAST, and/or Weka; or Ceph, LUSTRE, or similar open source tools.

Interview Process

After submitting your application, the team reviews resume. If your application passes this stage, you will be invited to a 15 minute hiring manager screen. If you clear the initial phone interview, you will enter the main process, which consists of three 45 minute interviews: a technical deep dive, customer communications and debugging session, and culture fit interview.

Our goal is to finish the main process within one week. All interviews will be conducted via virtually.

Benefits
  • Competitive total compensation package (cash + equity).

  • Retirement or pension plan, in line with local norms.

  • Health, dental, and vision insurance.

  • Generous PTO policy, in line with local norms.

  • Fluidstack is remote first, but has offices in London, New York, and SF. For all other locations, we provide access to WeWork.

Top Skills

Ansible
Bash
Go
Kubernetes
Python
Slurm
Terraform

Similar Jobs

2 Days Ago
In-Office or Remote
29 Locations
Senior level
Senior level
Information Technology • Web3
As a Site Reliability Engineer, you'll design scalable infrastructure, monitor performance, optimize services, lead incident responses, and collaborate across teams to enhance user experiences.
Top Skills: ArgocdErigonGethGithub ActionsGoogle Cloud BuildGoogle Cloud PlatformKafkaKubernetesPostgresPrometheusReth
15 Days Ago
Remote
28 Locations
Junior
Junior
Information Technology • Software • Travel • Hospitality
As a Site Reliability Engineer at Cloudbeds, you will ensure the reliability and performance of systems, support cloud infrastructure, improve monitoring, and collaborate across teams to optimize services.
Top Skills: ArgocdAuroraAWSBitbucketDatadogDockerGithub ActionsKubernetesLokiMemcachedMySQLNginxPostgresPrometheusRedisSqsTerraform
15 Days Ago
Remote
28 Locations
Junior
Junior
Information Technology • Software • Travel • Hospitality
As a Site Reliability Engineer at Cloudbeds, you will enhance the reliability and performance of systems, automate processes, and ensure optimal system operations in a remote team environment.
Top Skills: ArgocdAuroraAWSBitbucketDatadogDockerGithub ActionsKubernetesLokiMemcachedMySQLNginxPostgresPrometheusRedisSqsTerraform

What you need to know about the Dublin Tech Scene

From Bono and Oscar Wilde to today's tech leaders, Dublin has always attracted trailblazers, with more than 70,000 people working in the city's expanding digital sector. Continuing its legacy of drawing pioneers, the city is advancing rapidly. Ireland is now ranked as one of the top tech clusters in the region and the number one destination for digital companies, with the highest hiring intention of any region across all sectors.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account