Nebius
Jobs at Nebius
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Recently posted jobs
Artificial Intelligence • Information Technology • Consulting
Own the roadmap, reliability, scalability, and compliance of internal billing tools. Lead discovery with Finance, Support, Sales, and partners; translate billing, regulatory, and operational needs into product requirements. Drive automation for account setup, quota management, reconciliation, and billing corrections. Partner with Engineering, Finance, Legal, and Operations on RBAC/IAM, APIs, data pipelines, incident runbooks, SLAs, and OKRs that improve billing accuracy, auditability, onboarding, and operational efficiency.
Artificial Intelligence • Information Technology • Consulting
Serve as the primary technical advisor for strategic GPU cloud customers: troubleshoot complex AI/ML issues, optimize GPU performance, guide deployments at scale, collaborate with sales and product teams, and translate customer needs into product feedback and solutions.
Artificial Intelligence • Information Technology • Consulting
Own the design, delivery, and operation of complex data pipelines, datasets, and platform components. Translate business needs into scalable data models and reliable data products, improve data quality and observability, resolve production issues, and establish reusable engineering tools and standards. Collaborate with product and business stakeholders, contribute to architecture decisions, mentor engineers, and participate in on-call support.
Artificial Intelligence • Information Technology • Consulting
Develop and maintain API-based integrations and cloud-native Azure services. Build workflow automation, identity lifecycle processes, data synchronization, CI/CD pipelines, and Terraform infrastructure. Troubleshoot APIs, Azure services, and data sources while documenting integrations and scripts. The role requires Python or C#, Azure integration expertise, secure authentication knowledge, relational database experience, and familiarity with deployment automation.
Artificial Intelligence • Information Technology • Consulting
Drive cloud capacity expansion across data centers by analyzing resource consumption, forecasting demand, identifying risks, and translating forecasts into infrastructure requirements. Coordinate procurement, hardware delivery, installation, deployment, and customer availability across supply chain, data center, engineering, product, and business teams. Maintain expansion plans, communicate timelines and dependencies, and proactively resolve cross-functional issues.
Artificial Intelligence • Information Technology • Consulting
Own and evolve the People Systems roadmap, partnering with HR stakeholders and the HiBob vendor on system enhancements, documentation, automation, self-service, and AI-driven support. Coordinate HRIS activities for mergers and acquisitions, improve data and lifecycle processes, and support analytics, reporting, and continuous improvement. Serve as deputy to the People Systems Manager while collaborating with technical and non-technical stakeholders in a global environment.
Artificial Intelligence • Information Technology • Consulting
Lead mechanical due diligence and feasibility for data center sites and deals. Validate cooling topology and equipment strategy (air, liquid, hybrid), assess capacity, vendor lead times, and integration with budgets and schedules. Define mechanical readiness gates, deliver validated scope to Delivery, coordinate with stakeholders and vendors, and help build repeatable evaluation standards for AI/hyperscale infrastructure acquisitions.
Artificial Intelligence • Information Technology • Consulting
Lead connectivity lifecycle for data center and PoP sites: assess site feasibility, design fiber pathways and POE/MMR layouts, manage carrier relationships and commercial negotiations (IRUs, leases), maintain geospatial/fiber documentation (KMZ/KML), and coordinate cross-functional delivery with procurement, construction and engineering teams across the region.
Artificial Intelligence • Information Technology • Consulting
Lead and coordinate complex, cross-functional compute hardware and software programs (kernel, virtualization, drivers, firmware, GPUs). Manage dependencies, risks, execution, deployment automation, validation, observability, and customer/vendor interactions to scale large GPU/HPC cloud infrastructure.
Artificial Intelligence • Information Technology • Consulting
Lead applied research and system design for agent-native search: develop multi-stage retrieval, ranking, and grounding methods for LLMs working with real-time web data; define evaluation paradigms; optimize relevance, latency, and cost; deploy solutions to production and mentor the engineering team.
Artificial Intelligence • Information Technology • Consulting
Design, train, and deploy retrieval, reranking, and relevance ML models for a production agent-native search platform. Build embedding-based indexing and large-scale retrieval systems, define evaluation metrics and pipelines, optimize latency/quality/cost trade-offs, and collaborate with engineers to integrate models into 24x7 production services.
Artificial Intelligence • Information Technology • Consulting
Design, implement, and operate the Virtual Private Cloud control and data planes. Improve performance, reliability, automation, and cross-service interactions. Develop networking features, perform load testing, and write high-performance code for cloud networking services.
Artificial Intelligence • Information Technology • Consulting
Lead hands-on applied AI work: prototype demos, accelerate customer POC-to-production, perform technical onboarding, optimize distributed training and inference on GPU infrastructure, research and implement emerging ML techniques, and feed actionable product feedback while producing reusable assets and benchmarks.
Artificial Intelligence • Information Technology • Consulting
Lead benchmarking and profiling of GPU platforms for ML/AI workloads. Profile system and kernel-level GPU performance, debug and optimize ML training and inference, perform acceptance testing of GPU clusters, run experiments across GPU configurations and interconnects, and develop tooling and dashboards to visualize performance and bottlenecks.
Artificial Intelligence • Information Technology • Consulting
Lead and drive cross-team technical projects across engineering domains (incident, capacity, runtime infra). Align stakeholders including C-level, facilitate technical decisions, establish scalable processes, and collaborate with engineering, infrastructure, and business teams.
Artificial Intelligence • Information Technology • Consulting
Lead reliability and observability for compute nodes running VMs. Debug Linux user/kernel issues, troubleshoot CPU/memory/NUMA/cgroups, operate QEMU/KVM and container tech, design node-level metrics/logs/traces/SLIs/SLOs, run incident response and collaborate across platform, kernel, GPU and infrastructure teams.
Artificial Intelligence • Information Technology • Consulting
Work on YDB distributed storage and database systems to maximize performance on modern and legacy hardware (NVMe, HDD, DPUs). Reengineer components for low-latency, high-throughput workloads, profile and debug production systems, and design resilient distributed storage primitives. Collaborate on incident resolution and production deployments to support Nebius cloud services.
Artificial Intelligence • Information Technology • Consulting
Lead performance analysis and optimization of large-scale GPU/HPC clusters across hardware and software stacks. Troubleshoot real workloads, qualify and integrate new hardware, tune system configurations (networking, virtualization, kernel), support escalations, and collaborate with infrastructure teams and hardware vendors to ensure cluster performance meets expectations.
Artificial Intelligence • Information Technology • Consulting
Own product direction for Soperator, a Slurm-on-Kubernetes control plane for GPU clusters. Drive discovery, design, delivery and adoption; lead customer research; prioritize roadmap and metrics; coordinate across compute, storage, networking, observability and IAM teams; lead open-source strategy and community adoption.
Artificial Intelligence • Information Technology • Consulting
Maintain and scale DevTools SRE systems: improve user-facing developer workflows, build fault-tolerant self-healing architecture, optimize performance, modify open- and closed-source tools (GitLab, TeamCity), and support users while measuring improvements via metrics.