duvo.ai Logo

duvo.ai

Site Reliability Engineer (EU/UK Based - Remote)

Reposted 14 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in London, Greater London, England
Mid level
In-Office or Remote
Hiring Remotely in London, Greater London, England
Mid level
The Site Reliability Engineer will manage platform reliability and infrastructure, ensuring security and observability while automating deployments. Responsibilities include incident response, monitoring, and improving reliability practices.
The summary above was generated by AI
Who we are

Enterprise teams still copy data between systems all day. Work gets stuck in emails, legacy UIs, and handoffs. That chaos is costly, slow, and risky.

We're a fast-moving team on a mission to end it for good. Traction is strong and we're solving real problems for real customers—but to win, we need exceptional talent. We stay humble, do the work, and let results speak.

What we are building

We're building the AI operations platform for retail and CPG enterprises—a horizontal platform where AI agents execute end-to-end work across UIs and APIs with governance built in.

Where copilots stop, Duvo finishes the job. Business users specify the outcome; agents plan, act, request approvals on exceptions, and learn with every run. We start with a retail wedge (category management, supply chain, finance ops) where ROI is obvious, then expand to adjacent functions and sectors.

Velocity is our moat: ship fast, iterate faster, compound learning.

The role

You will own the reliability, security, and infrastructure that lets our platform run AI agents for enterprise customers. This isn't traditional web app SRE — our agents execute arbitrary code in sandboxes, make unpredictable external API calls, and run for hours. Keeping this reliable, secure, and observable is the job.

You'll be part of newly formed SRE team as one of the first teammembers. Infrastructure is currently owned collectively by product engineers — you'll take ownership, inherit real infrastructure (25+ Terraform modules, full OpenTelemetry pipeline, Prometheus/Grafana monitoring), and build the reliability practice from scratch.

Your unit of ownership: platform reliability, infrastructure, observability, and incident response. You own sandbox infrastructure and capacity; the AI Platform Engineer owns sandbox behavior and runtime logic.

We're a growing product team scaling into multiple initiatives, each with a lead, engineers, a design engineer, and an AI-focused engineer.

What we're looking for

These are non-negotiables—the things we'll specifically evaluate you on:

  • Distributed systems experience. You've designed and operated systems that scale. You understand failure modes, capacity planning, and the tradeoffs between consistency, availability, and latency in real production environments.

  • Security mindset. You'll handle enterprise data flowing through sandboxed environments, manage KMS encryption, configure Cloud Armor WAF rules, and ensure network isolation between tenant workloads. Security is a default consideration, not an afterthought.

  • Observability and incident response. You build monitoring and alerting that catches problems before customers do. When incidents happen, you lead structured responses, find root causes, and drive lasting fixes — not just restarts.

  • Infrastructure as code and automation. You automate everything you can. You've worked with IaC tools, CI/CD pipelines, and container orchestration in production. Manual runbooks make you uncomfortable.

  • Shipping and ownership. You don't just maintain systems — you improve them. You take ownership of reliability projects from proposal to production, and you measure the results.

  • Judgment on where to invest. You'll decide what to automate first, where to invest in reliability vs. ship speed, and make incident calls with incomplete information.

You might also
  • Have experience with GCP, Kubernetes, or similar cloud-native infrastructure.

  • Have worked with sandboxed execution environments or multi-tenant isolation.

  • Be comfortable with AI/ML production systems — understanding the unique reliability challenges of LLM-based applications.

  • Have a product engineering background — you've built features and understand the developer experience you're supporting.

This is not for you if
  • You want a traditional ops role where you follow runbooks — we're building the reliability practice, not maintaining one.

  • You want to build AI features — see AI Platform Engineer.

Our tech stack
  • GCP (Cloud Run, GKE, GCS)

  • Terraform, Docker

  • Prometheus, Grafana, Loki, OpenTelemetry

  • TypeScript and Python services (you'll read and occasionally modify application code, but deep language expertise isn't required)

  • Postgres, Redis

How we work

These are real tradeoffs we've made, not aspirations:

  • Initiative-driven. We organize around customer problems, not org charts. Problems surface through product feedback, competitive analysis, and direct customer conversations — then we prioritize, build, and ship weekly.

  • Customer-obsessed. We solve real problems, not hypothetical ones. Features that don't move customer metrics get cut.

  • Iterative by default. We ship small, learn fast, and never get attached to yesterday's code. This means things break sometimes — we fix forward.

  • AI-first leverage. We use AI to move faster and focus human time where it matters most. If a tool can do it, a person shouldn't.

  • Direct feedback. We give each other actionable feedback immediately. This can feel uncomfortable — we think that's worth it.

  • Autonomy with accountability. We trust people to make decisions and hold them to outcomes, not process.

What we offer
  • Unlimited AI budget. We don't just allow AI tools — we strongly encourage them. Want to try a new tool? Buy it. Want to automate part of your workflow? Do it.

  • Autonomy to do your best work. Want to meet someone to learn from? Set it up. Want a mentor? Go get one. Want to fly out to talk to an important customer? Just ask.

  • A real AI product with real customers. You're not building demos or internal tools. Enterprise customers use what you ship, and their feedback drives what you build next.

  • A sharp, motivated team that values ownership and candor.

  • Compensation 250.000,- CZK / month with a meaningful equity component. You can trade salary for additional equity if you prefer more upside.

How we hire

We respect your time and aim to move fast:

  1. Discovery call with a senior teammember (online, 30 min). We'll talk about you, how you think and whether there's mutual fit.

  2. Remote task (async, time-boxed, ~1 hour). Build a small product end-to-end. Not LeetCode.

  3. Technical interview (online, ~1 hour). Meet the team. We'll go deeper on your experience, system design, product thinking, and collaboration. No trick questions — we want to see how you think and build.

  4. On-site trial day (2 days). Ship something small to production with us and see how we work together. Fully compensated.

Similar Jobs

4 Hours Ago
In-Office or Remote
Entry level
Entry level
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Outbound, quota-carrying SDR role focused on Mid-Market DACH: prospecting, qualifying leads, cold calls/emails, converting opportunities into pipeline, and collaborating with sales, marketing, partners, and operations. Use Salesforce, Gong, Outreach, and LinkedIn Navigator to manage outreach and pipeline.
Top Skills: GongLinkedin NavigatorOutreachSalesforce
4 Hours Ago
In-Office or Remote
Junior
Junior
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Perform quota-carrying outbound prospecting for the Enterprise DACH market: cold calls, emails, and research to qualify leads and convert meetings into pipeline. Collaborate with sales, marketing, partners, and operations; handle objections with value messaging; write personalized outreach; build pipeline with Enterprise Advocates and Marketing; use Salesforce, Gong, Outreach, LinkedIn Navigator, and AI tools to improve research, execution, and analysis.
Top Skills: AIGongLinkedin NavigatorOutreachSalesforce
4 Hours Ago
In-Office or Remote
Junior
Junior
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Outbound, quota-carrying SDR focused on Enterprise Emerging Markets: prospecting, cold outreach, qualifying leads, converting meetings into pipeline, and partnering with sales, marketing and enterprise teams while using Salesforce, Gong, Outreach, LinkedIn Sales Navigator and AI tools.
Top Skills: AIGongLinkedin Sales NavigatorOutreachSalesforce

What you need to know about the Dublin Tech Scene

From Bono and Oscar Wilde to today's tech leaders, Dublin has always attracted trailblazers, with more than 70,000 people working in the city's expanding digital sector. Continuing its legacy of drawing pioneers, the city is advancing rapidly. Ireland is now ranked as one of the top tech clusters in the region and the number one destination for digital companies, with the highest hiring intention of any region across all sectors.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account