Level

HeyGen

Software Engineer, GPU Performance

AI in this role

Software engineer focused on GPU performance, optimization, and inference infrastructure for AI video applications.

pytorchpythoncudatritonnsight-systemsnsight-compute
gpu-optimizationprofilinginferenceperformance-engineering

About HeyGen

At HeyGen, our mission is to make visual storytelling accessible to all. Over the last decade, visual content has become the preferred method of information creation, consumption, and retention. But the ability to create such content, in particular videos, continues to be costly and challenging to scale. Our ambition is to build technology that equips more people with the power to reach, captivate, and inspire audiences.
Learn more at www.heygen.com.  Visit our Mission and Culture doc here. 

Position Summary

HeyGen is building AI applications including Avatar IV, Photo Avatar, Interactive Avatar, and Video Translation. We’re looking for a Software Engineer focused on GPU performance to make the systems behind these experiences faster and more efficient.

You will work across model execution and inference infrastructure, using profiling and measurement to improve latency, throughput, and GPU cost. This role is a fit for an engineer who enjoys understanding how software uses the hardware beneath it.

Key Responsibilities

  • Use NVIDIA Nsight Systems, Nsight Compute, and PyTorch Profiler to investigate GPU utilization, kernel execution, memory bandwidth, and CPU–GPU data movement.
  • Identify bottlenecks across model execution, preprocessing, and inference serving, then measure the impact of each optimization.
  • Improve performance through batching, scheduling, memory management, and better GPU utilization.
  • Develop or integrate high-performance GPU kernels when existing implementations limit performance.
  • Build benchmarks and automated checks that catch performance regressions across representative video workloads.
  • Collaborate with AI researchers and infrastructure engineers to bring optimizations into production.
  • Measure the effect of changes on latency, throughput, cost, and output quality.

Qualifications

  • Experience optimizing GPU-based AI workloads or high-performance computing systems.
  • Proficiency in Python and experience with PyTorch or a similar machine learning framework.
  • Strong curiosity about GPU hardware, including memory bandwidth, cache behavior, tensor cores, and data movement between CPU and GPU.
  • Experience using profiling tools to connect hardware behavior to application-level bottlenecks and validate improvements.
  • Ability to turn performance experiments into reliable production changes and communicate tradeoffs clearly.

Preferred Qualifications

  • Experience with CUDA, Triton, or C++ GPU programming.
  • Experience optimizing video, image, audio, diffusion, or Transformer models.
  • Familiarity with multi-GPU inference, GPU interconnects, quantization, or large-scale model serving.
  • Experience building performance benchmarks or regression testing infrastructure.
  • Prior experience in a fast-paced technology environment.

What HeyGen Offers

  • Competitive salary and benefits package.
  • Dynamic and inclusive work environment.
  • Opportunities for professional growth and advancement.
  • Collaborative culture that values innovation and creativity.
  • Access to the latest technologies and tools.

HeyGen is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Join us at HeyGen

Help us make AI video faster and more accessible. We’d love to hear from you.

How we rate this

Software Engineer, GPU Performance at HeyGen rates 85 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.

Classification

Builds AI. The job is building AI systems.

  1. ●●●● Builds AI80 to 100
  2. ●●●○ Works on AI60 to 79
  3. ●●○○ Uses AI40 to 59
  4. ●○○○ Little AI0 to 39

Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

GPU OptimizationProfilingInferencePerformance EngineeringPyTorchPythonCudaTriton

Questions you could be asked

  1. Tell me about a project where gpu optimization was part of your work. What did you do?
  2. Tell me about a project where profiling was part of your work. What did you do?
  3. Tell me about a project where inference was part of your work. What did you do?
  4. Tell me about a project where performance engineering was part of your work. What did you do?
  5. Walk me through how you've used PyTorch in your day-to-day work.

Adapt your resume

  • List these exact terms on your resume: GPU Optimization, Profiling, Inference, Performance Engineering, and PyTorch. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Get new AI jobs (Builds AI ●●●●) by email

One email a week with the new AI jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Software Engineering roles that build AI, at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at HeyGen

Related searches

Same AI level