Senior Research Engineer - Enterprise Products

NVIDIA Corporation
United States
4 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$192,000.0 - $304,750.0
Working hours
Regular working hours

Tech stack

Artificial Neural Networks Nvidia CUDA Distributed Systems Machine Learning Natural Language Processing Tensorflow System Software Speech Recognition Graphics Processing Unit (GPU) Pytorch Large Language Models Deep Learning
+2 more
Gpu Programming Information Technology

Requirements

  • Bachelor’s of Master’s degree in Computer Science or equivalent experience.
  • 8+ years of industry experience in Deep Learning frameworks (PyTorch or TensorFlow).
  • Experience designing or running LLM evaluations/benchmarks - ideally agentic ones - and drawing statistically sound conclusions from them
  • Understanding of modern techniques in Machine Learning, Deep Neural Networks, Natural Language Processing, or Speech Recognition.
  • Empirical research mindset: forming hypotheses about new algorithms, running calibrations, iterating on results
  • Strong communication and interpersonal skills, along with the ability to work in a dynamic and distributed team. A history of mentoring junior engineers and interns is a huge plus.
  • A desire to constantly grow and learn new things.
  • Strong computer science fundamentals - algorithms and data structures, computational complexity, parallel and distributed computing, system software.

Ways to stand out from a crowd:

  • Experience architecting or developing large-scale distributed systems for deep learning.
  • Agentic benchmark creation and publications.
  • Knowledge of CPU and/or GPU architecture.
  • GPU programming (CUDA).

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 192,000 USD - 304,750 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

About the company

We are now looking for a Senior Research Engineer passionate about Generative AI inference. Are you excited to change the way people infuse AI into products and services? NVIDIA is at the forefront of generative AI models, from language to images. NVIDIA provides building blocks to democratize AI and make generative AI easy to develop, integrate, and deploy. Our team is dedicated to developing optimized inferencing technologies to support our growing generative AI needs. We contribute to all steps of the machine learning lifecycle: from conceptualization, to applied research, engineering for optimized inference, and deployment. Collaborate with research teams, engineers, and open-source community.

What you will be doing:

  • Design and evaluate routing policies for LLM traffic to best use mixture of model systems.
  • Build and run agentic benchmarks (e.g., Terminal-Bench ) to measure algorithm quality, and turn results into calibration data and routing profiles
  • Ship to an open-source repo: design docs, code review, docs, and community contributions
  • Collaborating with engineering teams across all of NVIDIA to ensure our software integrates seamlessly up and down the NVIDIA accelerated serving stack.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

2:08 min

History and scale of NVIDIA GPU computing

Paul Graham Paul Graham · LIVE

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

Videos

See all

Related articles

See all