Nebius
Jobs at Nebius
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Recently posted jobs
Artificial Intelligence • Information Technology • Consulting
Lead and build Nebius' Detection & Response capability: design detections, achieve MITRE coverage, run incident response and forensics, integrate threat intelligence, build D&R tooling and SOAR workflows, and manage cross-functional stakeholders and metrics (MTTD/MTTR). Hire, mentor, and scale a team and on-call program.
Artificial Intelligence • Information Technology • Consulting
Build and operate the network platform: set SLIs/SLOs, drive reliability improvements, own incident response and postmortems, develop observability and safer change workflows, automate operations, and collaborate with network and platform teams to embed operability.
Artificial Intelligence • Information Technology • Consulting
Build and maintain services and tooling to automate and secure network lifecycle across data centers. Implement CI/CD, staged rollouts, drift detection, observability pipelines, and APIs that integrate network source-of-truth with the cloud platform. Collaborate with network engineers and SREs to turn operational pain into production-grade automation and tooling.
Artificial Intelligence • Information Technology • Consulting
As a Senior Support Engineer, you will troubleshoot complex issues in cloud environments, focusing on Linux and Kubernetes, and assist customers with AI workloads. You will also improve internal tools and processes while communicating effectively with clients during incidents.
Artificial Intelligence • Information Technology • Consulting
Support mechanical and electrical technical evaluations for data center site acquisitions: run capacity calculations, compare equipment, assess cooling and power distribution, track lead times and risks, prepare Go/No-Go materials, and coordinate with vendors and senior engineers to hand off validated scope to Delivery.
Artificial Intelligence • Information Technology • Consulting
Lead mechanical due diligence and feasibility for data center sites and deals. Validate cooling topology and equipment strategy (air, liquid, hybrid), assess capacity, vendor lead times, and integration with budgets and schedules. Define mechanical readiness gates, deliver validated scope to Delivery, coordinate with stakeholders and vendors, and help build repeatable evaluation standards for AI/hyperscale infrastructure acquisitions.
Artificial Intelligence • Information Technology • Consulting
Lead connectivity lifecycle for data center and PoP sites: assess site feasibility, design fiber pathways and POE/MMR layouts, manage carrier relationships and commercial negotiations (IRUs, leases), maintain geospatial/fiber documentation (KMZ/KML), and coordinate cross-functional delivery with procurement, construction and engineering teams across the region.
Artificial Intelligence • Information Technology • Consulting
Lead and coordinate complex, cross-functional compute hardware and software programs (kernel, virtualization, drivers, firmware, GPUs). Manage dependencies, risks, execution, deployment automation, validation, observability, and customer/vendor interactions to scale large GPU/HPC cloud infrastructure.
Artificial Intelligence • Information Technology • Consulting
Design, operate and troubleshoot large-scale data center and backbone networks including InfiniBand GPU cluster interconnects. Provide technical design and operational support, develop monitoring and automation tools, document network designs, test infrastructure and vendors, and support launch of new cloud regions.
Artificial Intelligence • Information Technology • Consulting
Lead applied research and system design for agent-native search: develop multi-stage retrieval, ranking, and grounding methods for LLMs working with real-time web data; define evaluation paradigms; optimize relevance, latency, and cost; deploy solutions to production and mentor the engineering team.
Artificial Intelligence • Information Technology • Consulting
Design, train, and deploy retrieval, reranking, and relevance ML models for a production agent-native search platform. Build embedding-based indexing and large-scale retrieval systems, define evaluation metrics and pipelines, optimize latency/quality/cost trade-offs, and collaborate with engineers to integrate models into 24x7 production services.
Artificial Intelligence • Information Technology • Consulting
Design, implement, and operate the Virtual Private Cloud control and data planes. Improve performance, reliability, automation, and cross-service interactions. Develop networking features, perform load testing, and write high-performance code for cloud networking services.
Artificial Intelligence • Information Technology • Consulting
Design, deploy, and optimize global enterprise wired/wireless networks and VPNs. Lead connectivity between branches, automate operations and monitoring, troubleshoot cross-region incidents, evaluate new technologies, and partner with telecom providers to manage SLAs and escalations.
Artificial Intelligence • Information Technology • Consulting
Lead hands-on applied AI work: prototype demos, accelerate customer POC-to-production, perform technical onboarding, optimize distributed training and inference on GPU infrastructure, research and implement emerging ML techniques, and feed actionable product feedback while producing reusable assets and benchmarks.
Artificial Intelligence • Information Technology • Consulting
Own product vision, roadmap, and backlog for WAN and underlay networking (interconnects, backbone, data-center fabric). Define Direct Connect-style products, throughput tiers, QoS and traffic engineering. Coordinate capacity planning, cross-region and cross-cloud connectivity, and SLAs. Work closely with network engineering, product marketing, pre/post-sales, and support to deliver reliable, high-throughput WAN services.
Artificial Intelligence • Information Technology • Consulting
Lead reliability and observability for compute nodes running VMs. Debug Linux user/kernel issues, troubleshoot CPU/memory/NUMA/cgroups, operate QEMU/KVM and container tech, design node-level metrics/logs/traces/SLIs/SLOs, run incident response and collaborate across platform, kernel, GPU and infrastructure teams.
Artificial Intelligence • Information Technology • Consulting
Lead and drive cross-team technical projects across engineering domains (incident, capacity, runtime infra). Align stakeholders including C-level, facilitate technical decisions, establish scalable processes, and collaborate with engineering, infrastructure, and business teams.
Artificial Intelligence • Information Technology • Consulting
Design and implement cloud infrastructure and MLOps solutions for ML/AI customers. Advise clients, run PoCs and workshops, create IaC and technical documentation, optimize GPU training/inference pipelines, and act as primary technical liaison across product, support, and marketing teams.
Artificial Intelligence • Information Technology • Consulting
Serve as the primary technical advisor for strategic GPU cloud customers: troubleshoot complex AI/ML issues, optimize GPU performance, guide deployments at scale, collaborate with sales and product teams, and translate customer needs into product feedback and solutions.
Artificial Intelligence • Information Technology • Consulting
Administer and improve the Atlassian Cloud ecosystem (Jira, Confluence, Assets, Statuspage), design workflows and automations, manage SSO/SCIM with Azure Entra ID, integrate via APIs/webhooks, ensure SaaS governance and security compliance, monitor performance, document configurations, and provide user support and training across enterprise teams.