Optimize large-model inference at scale: improve throughput, reduce latency and cost per token, build benchmarking harnesses, tune parallelism and quantization strategies, implement load-balancing in routing, and evaluate custom kernels and emerging inference hardware.
Dragonfly is a crypto-native Venture Capital and research firm with $3.6B+ in assets under management and 160+ portfolio companies. Our Talent team connects people with roles across our portfolio, opening the door to opportunities through our Talent Network.
This is an application to join our talent network. This is not a listing for an internal role at Dragonfly.
We're actively sourcing for a Senior Inference Optimization Engineer for one of our portfolio companies building privacy-first consumer AI infrastructure. You'll be on the bleeding edge of LLM inference performance, pushing throughput, driving down latency, and optimizing cost per token at significant scale.
Location: Remote, USA (open to excellent candidates outside the USA)
What We’re Looking For:
- 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
- Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume
- Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile
- Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
- GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
- Proficiency in Python, Rust, or Go. C++/CUDA a strong plus
- Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions
About the role:
- Stand up and optimize GPU infrastructure including B300 nodes in owned data centers
- Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads
- Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
- Optimize multivariate inference load-balancing algorithms within the inference routing system
- Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements
- Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack
Even if you don't match every point above but are an engineer passionate about AI and/or crypto, we encourage you to apply. There may be other opportunities that fit your skill set.
Process:
- We'll review your application and assess fit for this role.
- If there's a match, we'll facilitate a warm introduction to the team.
- If the timing isn't right, we'll keep you in mind for future opportunities across the portfolio.
Submit your information below, and we’ll reach out if there’s a potential fit.
Similar Jobs
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Leads strategy, delivery, governance, adoption, and lifecycle management for agentic AI products across Pfizer’s regulated manufacturing network. Defines use cases, roadmaps, evaluation frameworks, guardrails, success metrics, and value-realization models while partnering with manufacturing, engineering, data, quality, cybersecurity, and business teams. Oversees pilots, production deployment, monitoring, continuous improvement, responsible AI, and scaling of intelligent manufacturing capabilities.
Top Skills:
Agentic AiAICloud Data PlatformsHistoriansLimsLlmopsMesMlopsPlant ConnectivityScadaSerialization
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Build, prototype, evaluate, deploy, and improve agentic AI products for Pfizer’s regulated manufacturing network. Own agent backlogs, user discovery, experimentation, adoption, success metrics, governance, responsible AI, monitoring, and lifecycle operations. Partner with engineers, data scientists, manufacturing SMEs, Quality, Cybersecurity, and business stakeholders to convert manufacturing knowledge and user needs into production-grade intelligent agents that improve productivity, decision-making, and operational performance.
Top Skills:
Advanced AnalyticsAgentic AiAgentic Ai OrchestrationAgile/ScrumAi Agent FrameworksAi Evaluation FrameworksAi GuardrailsCloud Data PlatformsCsvHistoriansHuman-In-The-Loop OversightLarge Language Models (Llms)LimsLlmopsMesMlopsOt/AutomationPrompt EngineeringRetrieval-Augmented Generation (Rag)Scada
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead design, delivery and maintenance of GMP/GDP audit strategy for biologics, aseptic, small molecule and medical device areas. Plan and execute complex audits, coach auditors, analyze regulatory intelligence, drive corrective actions, support inspection readiness, and partner with PGS/PharmSci and site stakeholders to ensure compliance and continuous improvement.
Top Skills:
Aseptic ManufacturingDigital HealthEu DirectiveFdaGdpGmpIchIsoMedical Device RegulationsPic/SQuality Management SystemSoftware As A Medical Device (Samd)Tga
What you need to know about the Dublin Tech Scene
From Bono and Oscar Wilde to today's tech leaders, Dublin has always attracted trailblazers, with more than 70,000 people working in the city's expanding digital sector. Continuing its legacy of drawing pioneers, the city is advancing rapidly. Ireland is now ranked as one of the top tech clusters in the region and the number one destination for digital companies, with the highest hiring intention of any region across all sectors.

