eBay Logo

eBay

SRE- Availability Engineer

Posted 8 Days Ago
Be an Early Applicant
In-Office
Dublin, IRL
Senior level
In-Office
Dublin, IRL
Senior level
Lead incident response for critical production systems, triage and coordinate multi-team mitigation, improve availability through monitoring, automation, and post-incident improvements, and act as senior technical escalation on-shift.
The summary above was generated by AI

At eBay, we're more than a global ecommerce leader — we’re changing the way the world shops and sells. Our platform empowers millions of buyers and sellers in more than 190 markets around the world. We’re committed to pushing boundaries and leaving our mark as we reinvent the future of ecommerce for enthusiasts.

Our customers are our compass, authenticity thrives, bold ideas are welcome, and everyone can bring their unique selves to work — every day. We're in this together, sustaining the future of our customers, our company, and our planet.

Join a team of passionate thinkers, innovators, and dreamers — and help us connect people and build communities to create economic opportunity for all.

Looking for a company that inspires passion, courage, and creativity, where you can help shape the future of global commerce? At eBay, millions of people around the world come to buy, sell, connect, and share. We are building a more ambitious, inclusive, and customer-focused future — and we’re looking for people who want to help power it.

The opportunity

eBay is seeking an experienced Senior Site Reliability Engineer (Technical Duty Officer ) to help protect the availability, stability, and resilience of our most critical production systems. This is a high-impact role for a senior technical operator who thrives under pressure, leads calmly during incidents, and can quickly turn incomplete signals into conclusive action.

The Technical Duty Officer (TDO) is the senior technical leader within the shift for active production events in the Site Engineering Center. When unexpected outages or degradations occur, the TDO rapidly assesses technical and business impact, drives incident triage, coordinates multi-functional responders, and leads mitigation through restoration.. When unexpected outages or degradations occur, the TDO rapidly assesses technical and business impact, drives incident triage, coordinates multi-functional responders, and leads mitigation through restoration. The role advances the Availability pillar of Site Reliability at eBay. It provides senior technical judgment and strengthens operational standards. It also improves response rigor and resilience practices to boost system availability. The Site Engineering Center is part of eBay’s Site Reliability Engineering organization, and the TDO works closely with other SRE pillars, including Reliability, Observability, and Performance, to drive strong reliability outcomes across the platform. The role sits at the center of eBay’s incident management and service reliability strategy, partnering across engineering, architecture, SRE, operations, and support teams to improve availability at both the incident level and the systems level.

This role is ideal for someone with broad technical depth, strong operational judgment, and the ability to lead across infrastructure, application, database, network, and platform boundaries. The strongest candidates bring a customer-first demeanor, clear communication, and the discipline to improve systems not just during incidents, but after them as well.

What you’ll do
  • Lead the technical response during high-priority incidents, including assessing severity, clarifying customer and business impact, setting direction, coordinating responders, and driving restoration.

  • Maintain awareness of the health of eBay’s critical services and key business indicators, identifying risks early and driving action before they impact customers or the business.

  • Serve as the senior technical escalation point on shift, using sound judgment to guide mitigation, recovery, and prioritization during complex live-site events.

  • Partner across engineering, architecture, SRE, and operations to improve availability through stronger recovery strategies, monitoring, alerting, automation, and operational tooling.

  • Drive operational excellence through clear incident documentation, stakeholder communications, emergency change support, maintenance oversight, and continuous improvement activities such as drills, audits, and post-incident follow-through.

What you’ll bring
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience, along with 5+ years in site reliability engineering, production operations, incident management, systems administration, or infrastructure engineering.

  • Proven experience supporting highly scaled internet or platform environments, with broad troubleshooting capability across infrastructure, compute, containers, networking, storage, databases, DNS, load balancing, observability, and high availability concepts.

  • Proficiency in creating software solutions for task automation using technologies such as Go, Python, Java, Node.js, Docker, and Kubernetes to improve operational efficiency and reliability.

  • Strong incident leadership and operational judgment, including the ability to triage quickly, prioritize effectively, separate symptoms from causes, and make sound decisions under pressure.

  • Superb communication and influence skills, with the ability to lead across organizational boundaries and communicate clearly with engineers, senior leaders, and non-technical partners during high-stakes events.

  • A strong ownership attitude and calm execution style.

Preferred qualifications
  • Experience serving as an incident commander, major incident manager, or senior on-shift technical lead

  • Experience in e-commerce, large-scale consumer platforms, or other critically important online environments

  • Familiarity with executive communications during production incidents

  • Experience improving monitoring, alert quality, response automation, or postmortem effectiveness

  • Comfort using data, telemetry, dashboards, and change history to form and validate technical hypotheses during live incidents

Why this role matters

The Site Engineering Center exists to protect the community and internal customers by ensuring the reliability of critical services through strong detection, repair, and continuous improvement. As a Technical Duty Officer, you will sit at the center of that mission. Your leadership will directly influence site stability, customer trust, and the speed and quality of eBay’s response when the unexpected happens.

Shift and location details

This role is part of the Site Engineering Center and operates on a fixed shift schedule. This position is based in Dublin, Ireland and is aligned to a day shift. Team members work four consecutive 10-hour shifts, including one weekend day. This is a true on-shift operational role — you will not be on call outside your scheduled hours.

Additional Details

eBay is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex, sexual orientation, gender identity, veteran status, and disability, or other legally protected status. If you have a need that requires accommodation, please contact us at [email protected]. We will make every effort to respond to your request for accommodation as soon as possible. View our accessibility statement to learn more about eBay's commitment to ensuring digital accessibility for people with disabilities.


We use cookies to enhance your experience and may use AI tools for administrative tasks in the hiring process. To learn how we handle your personal data and use AI responsibly, please visit our Talent Privacy Notice, Privacy Center, and AI Hiring Guidelines.

eBay Dublin, Dublin, IRL Office

Dublin, Ireland

Similar Jobs

18 Minutes Ago
Easy Apply
Hybrid
Dublin, IRL
Easy Apply
Entry level
Entry level
Big Data • Cloud • Software • Database
The Account Development Representative will identify potential clients, generate leads, and support the sales team in achieving sales goals. Responsibilities include managing leads in Salesforce and developing product knowledge.
Top Skills: Salesforce
An Hour Ago
Hybrid
2 Locations
Mid level
Mid level
Financial Services
Design, build, and operate enterprise-scale infrastructure platforms. Write and review production code, diagnose production issues, participate in on-call rotations, apply AI-assisted development, and improve scalability, reliability, security, and cost. Collaborate with stakeholders, decompose problems, measure impact, and contribute to platform-level technical decisions in a greenfield Dublin engineering hub.
Top Skills: Ai ToolsC++Ci/CdDatabasesGoJavaKubernetesLinuxNetworkingPythonRustTerraformUnix
2 Hours Ago
Remote or Hybrid
2 Locations
Senior level
Senior level
Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Lead product vision and roadmap for agentic AI solutions, own the backlog, translate business needs into user stories, partner with data scientists and engineers, manage Agile delivery cycles, ensure responsible AI practices, and monitor model performance and continuous improvement.
Top Skills: Agentic AiAPIsCloud ServicesConfluenceJIRALlmsMachine LearningMicroservicesModel Evaluation/MonitoringModel Training/Fine-TuningRestSoapVersion One

What you need to know about the Dublin Tech Scene

From Bono and Oscar Wilde to today's tech leaders, Dublin has always attracted trailblazers, with more than 70,000 people working in the city's expanding digital sector. Continuing its legacy of drawing pioneers, the city is advancing rapidly. Ireland is now ranked as one of the top tech clusters in the region and the number one destination for digital companies, with the highest hiring intention of any region across all sectors.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account