Level

NVIDIAPosted 1w ago

Senior Deep Learning Solution Architect

Senior Deep Learning Solution Architect at NVIDIA scores 90 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.

China, ShanghaiseniorFull time

AI in this role

Senior Deep Learning Solution Architect to optimize LLM inference, training acceleration, and develop open-source inference frameworks.

vllmsglangflexkv
deep-learningdistributed-trainingperformance-optimizationllm-inference

NVIDIA is leading company of AI computing. At NVIDIA, our employees are passionate about AI, HPC , VISUAL, GAMING. Our SA team is more focusing to bring NVIDIA new technology into difference industries. We help to design the architecture of AI computing platform, analysis the AI and HPC applications to deliver our value to customers, focusing on defining and solving computational challenges in LLM inference and training acceleration, as well as network communication and data transfer optimization.

What You'll Be Doing:

  • Contribute to the development of open-source inference frameworks such as SGLang and vLLM, including feature and operator development, performance optimization, and model support, in collaboration with the community.

  • Develop and optimize KV cache offloading frameworks for LLM workloads, supporting multi-level cache offloading and reuse across CPU, SSD, and remote storage to improve inference efficiency. (Team project: FlexKV)

  • Drive R&D on compute performance in distributed training, and explore methods and technologies for performance optimization.

  • Study computational challenges in machine learning systems, identify common needs and bottlenecks, and build example code, acceleration libraries, or frameworks accordingly.

What We Need to See:

  • Over 5 years working experience in the technology industry, with master’s degree or above in computer science, mathematics, electrical engineering, automation, or related fields.

  • Strong interest in accelerated computing, parallel computing, and heterogeneous computing, with the motivation to explore these areas in depth.

  • Solid programming skills, with a good understanding of data structures and computer systems fundamentals.

  • Strong learning agility, adaptability, and the ability to analyze, define, and independently explore technical problems.

Ways to Stand Out from the Crowd:

  • Familiarity with heterogeneous computing, distributed training, parallel computing, or other areas related to high-performance computing.

  • Experience in performance analysis, performance modeling, or performance optimization; contributions to open-source frameworks are a plus.

  • Strong ability to define new problems and explore solutions; candidates with independent PhD-level research experience are preferred.

  • Proficiency with AI coding tools.

With competitive salaries and a generous benefits package, we are widely considered to be one of the world’s most desirable employers! We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our best-in-class engineering teams are rapidly growing. If you're a creative and autonomous person with a real passion for technology, we want to hear from you.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Deep LearningDistributed TrainingPerformance OptimizationLlm InferencevLLMSglangFlexkv

Questions you could be asked

  1. Tell me about a project where deep learning was part of your work. What did you do?
  2. Tell me about a project where distributed training was part of your work. What did you do?
  3. Tell me about a project where performance optimization was part of your work. What did you do?
  4. Tell me about a project where llm inference was part of your work. What did you do?
  5. Walk me through how you've used vLLM in your day-to-day work.

Adapt your resume

  • List these exact terms on your resume: Deep Learning, Distributed Training, Performance Optimization, Llm Inference, and vLLM. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Similar roles

Software Engineering roles rated AI Level 4 at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at NVIDIA

More solutions architect jobs

Related searches

Same AI level