Level

NVIDIAPosted 1d ago

L4

AI Computing Software Development Intern - 2027

AI Computing Software Development Intern - 2027 at NVIDIA scores 90 out of 100 on AI centrality, which makes it a Level 4 role on this board.

China, ShanghaiinternFull time

AI in this role

Build and enhance high-performance LLM inference pipelines, compiler graph optimizations, and GPU compute kernels for AI workloads.

pytorchpythonccudatensorrtllvmmlir
inference-optimizationcompiler-optimizationgpu-computingdeep-learning

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

 

What You’ll be doing:

As an intern, you’ll focus on one of three specialized tracks:

  • TensorRT LLM – Inference Optimization (Python / PyTorch): Build and enhance high‑performance LLM inference pipelines. Analyze and optimize model execution, scalability, and memory usage. Collaborate across framework and research teams to deliver efficient multi‑GPU model serving.

  • TensorRT Compiler – Graph Optimization (C++): Work on the TensorRT compiler backend to improve graph transformations and code generation for NVIDIA GPUs. Develop compiler optimization passes, refine operator fusion and memory allocation, and collaborate with CUDA and hardware architecture teams.

  • CuTe DSL & CUDA Kernels Development / Optimization (CUDA C++ / Python DSL): Design and tune GPU compute kernels and DSL implementations for core deep learning operations such as GEMM, MoE, Attention and Convolution. Profile, analyze and improve CUDA kernel performance to achieve maximum GPU efficiency.

 

What We need to see:

  • Pursuing an M.S. or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or related field.

  • Excellent problem‑solving ability, curiosity for cutting‑edge AI systems, and passion for GPU computing and deep learning software performance.

  • TensorRT LLM: Strong Python programming and experience with PyTorch; solid understanding of inference and GPU acceleration.

  • TensorRT Compiler: Proficient in C++, with experience in compiler or performance optimization.

  • CuTe DSL & CUDA Kernels: Skilled in C/C++ and CUDA or parallel programming; familiar with LLVM, MLIR and compiler; understanding of computer architecture and performance profiling/analysis/optimization.

 

Join us and play a part in building the AI computing platforms that drive innovation across industries worldwide.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Inference OptimizationCompiler OptimizationGpu ComputingDeep LearningPyTorchPythonCCuda

Questions you could be asked

  1. Tell me about a project where inference optimization was part of your work. What did you do?
  2. Tell me about a project where compiler optimization was part of your work. What did you do?
  3. Tell me about a project where gpu computing was part of your work. What did you do?
  4. Tell me about a project where deep learning was part of your work. What did you do?
  5. Walk me through how you've used PyTorch in your day-to-day work.

Adapt your resume

  • List these exact terms on your resume: Inference Optimization, Compiler Optimization, Gpu Computing, Deep Learning, and PyTorch. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Similar roles

Other roles rated Level 4 at other companies.

More jobs at NVIDIA

Related searches

Same AI level