Senior Site Reliability Engineer

Posted 2 Days Ago
Be an Early Applicant
Dublin
Senior level
Information Technology • Mobile • News + Entertainment • Social Media
The Role
As a Senior Site Reliability Engineer at Reddit, you will enhance the reliability and performance of engineering platforms. You'll collaborate with engineering teams to build resilient systems, automate repetitive tasks, diagnose and fix issues, and contribute to open-source projects, focusing on improving operational excellence.
Summary Generated by Built In

Reddit is a community of communities. It’s built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 97M+ daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit redditinc.com.

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities:

  • Advise
    • Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.
  • Amplify
    • Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit. 
    • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
    • Identify and engineer away risk across Reddit’s systems.
  • Automate
    • Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
    • Automate critical aspects of the event driven development process
  • Diagnose
    • Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.
    • Share on-call responsibilities. 
  • Optimize:
    • Observe and improve performance, reduce cost, and improve the experience for millions of users
    • Contribute upstream changes to the open source projects we use

Qualifications

  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience with the development and operation of high-traffic backend systems.
  • A demonstrated ability to debug, fix, and optimize code.
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.

Benefits:

  • Private Medical, Dental and Vision Benefits 
  • Retirement Savings plan with matching contributions
  • Workspace benefits for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Commuter Benefits  
  • Flexible Vacation & Reddit Global Days Off


Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve.  Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please contact us at [email protected].

Top Skills

Go
Python
The Company
HQ: San Francisco, CA
1,900 Employees
Hybrid Workplace
Year Founded: 2005

What We Do

Reddit is a community of millions of users engaging in the creation of content and the sharing of conversation across tens of thousands of topics. Our mission is to bring community, belonging, and empowerment to everyone in the world.

Why Work With Us

At Reddit, you’ll help build something that encourages millions around the world to think more, do more, learn more, feel more– and maybe even laugh more.

Gallery

Gallery

Similar Jobs

Hybrid
Dublin, IRL
289097 Employees

Klaviyo Logo Klaviyo

Sales Manager, SMB - Central Europe

Consumer Web • eCommerce • Marketing Tech • Retail • Software • Analytics • Generative AI
Hybrid
Dublin, IRL
2000 Employees

ServiceNow Logo ServiceNow

Senior Reliability Engineer - 24/7 Rotation Shifts

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Dublin, IRL
26000 Employees

Squarespace Logo Squarespace

Engineering Manager, Incident Response & Analysis

Consumer Web • eCommerce • Marketing Tech • Payments • Software • Design • SEO
Dublin, IRL
1723 Employees

Similar Companies Hiring

Dynatrace Thumbnail
Software • Information Technology • Cloud • Big Data Analytics • Big Data • Automation • Artificial Intelligence
Waltham , MA
4700 Employees
Citadel Thumbnail
Software • Information Technology • Financial Services • Big Data Analytics
Miami, FL
4000 Employees
Citadel Securities Thumbnail
Software • Information Technology • Financial Services
Miami, FL
1900 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account