Level

Hume AIPosted today

L4

Staff Software Engineer - Inference Backends

Staff Software Engineer - Inference Backends at Hume AI scores 97 out of 100 on AI centrality, which makes it a Level 4 role on this board.

Hume AI HQleadFullTime$170k-$280k

AI in this role

vllmpytorch

Hume AI is looking for a systems-oriented engineer to own the path from trained model to production inference. Join us in the heart of New York City and contribute to our endeavor to ensure that AI is guided by human values, the most pivotal challenge—and opportunity—of the 21st century.

About Us

Hume AI is a Series B startup dedicated to building artificial intelligence that is directly optimized for human well-being. As the first company to release speech language models, we’re focused on expanding our research to encompass audio understanding models and evaluation platforms for enterprises.


Our goal is to enable a future in which technology draws on an understanding of human emotional expression to better serve human goals. As part of our mission, we also conduct groundbreaking scientific research, publish in leading scientific journals like Nature, and support a non-profit, The Hume Initiative, that has released the first concrete ethical guidelines for empathic AI (www.thehumeinitiative.org). You can learn more about us on our website (https://hume.ai/) and read about us in WIRED, Forbes, and Venturebeat.

About the Role

As the first engineer dedicated full-time to inference backends at Hume, you will own the systems that take trained models from checkpoints to production inference: graph export, engine compilation, runtime integration, serving contracts, client libraries, artifact verification, and the performance and correctness of what runs in production.

You will work closely with research scientists, machine learning engineers, backend engineers, and the Data Plane team to bring new models and inference capabilities into production.

This is a systems-oriented role focused on performance, numerical correctness, reliability, and efficient use of accelerators. You will operate with a high degree of autonomy, make sound architectural decisions, and own systems throughout their lifecycle—from initial design and implementation through deployment, observability, optimization, and production support.

About the Role

As the first engineer dedicated full-time to inference backends at Hume, you will own the systems that take trained models from checkpoints to production inference. This includes graph export, engine compilation, runtime integration, serving contracts, artifact verification, and production performance and correctness.

You will also own Hume’s internal inference platform, which serves many of our in-house models across products. This includes deployment, routing, load balancing, health checking, observability, capacity management, and safe model rollout.

You will work closely with research scientists, machine learning engineers, backend engineers, and product teams to turn new model capabilities into reliable production systems.

This is a systems-oriented role focused on performance, numerical correctness, reliability, and efficient use of accelerators. You will operate with a high degree of autonomy and own systems from design through production.

What You’ll Do

  • Own the path from trained checkpoint to served request, including graph export, engine compilation, runtime integration, and serving configuration.

  • Build and evolve Hume’s internal inference platform for serving multiple models across products and workloads.

  • Design and operate serving infrastructure, including routing, load balancing, health checking, autoscaling, capacity management, and failure handling.

  • Build reproducible, versioned inference artifacts and tooling for validation, deployment, promotion, and rollback.

  • Build verification gates that catch numerical, behavioral, and performance regressions before production.

  • Design and maintain internal client libraries and standardized serving contracts.

  • Profile and optimize latency, throughput, memory usage, batching, scheduling, and accelerator utilization.

  • Diagnose production issues across application, runtime, container, networking, driver, and hardware boundaries.

  • Improve observability, resilience, testability, and operational safety across the inference stack.

  • Write clear technical documentation for the systems and APIs you build.

What You’ll Bring

  • Significant professional experience building server-side, infrastructure, distributed, or systems software.

  • Strong Linux fundamentals and hands-on experience troubleshooting and profiling production systems.

  • Professional experience with at least one systems-oriented language such as Rust, Go, C, or C++.

  • Experience with distributed systems concepts such as load balancing, health checking, failure recovery, observability, and capacity management.

  • A practical understanding of neural-network execution, including computation graphs, tensor shapes, data types, and accelerator execution.

  • Experience profiling and optimizing production systems.

  • Comfort working across languages and tooling, including Python for model export, validation, and integration workflows.

  • Strong ownership, independent technical judgment, and clear written and verbal communication.

  • The ability to use AI-assisted coding tools effectively while retaining the ability to explain, validate, debug, and modify the result independently.

Bonus Points

  • Experience building or operating shared model-serving or inference platforms.

  • Experience with GPU inference, CUDA, ONNX, TensorRT, PyTorch, Triton, vLLM, or similar systems.

  • Experience diagnosing numerical correctness issues such as precision loss, numerical drift, or nondeterminism.

  • Experience building high-performance client libraries, SDKs, or networked systems.

  • Experience with containers, CI/CD, or deploying software into on-premises or customer-managed environments.

  • Contributions to systems, inference, distributed-systems, or machine-learning infrastructure open-source projects.

L4Liked this Level 4 role? Get the best new ones weekly

Level 4 means “AI is the job”. Every week we send the highest-scoring new roles, Level 4 included, each rated Level 1 to 4. One email, no recruiter spam.

Free. One useful digest a week. Unsubscribe in one click.

Similar roles

Software Engineering roles rated Level 4 at other companies.

More jobs at Hume AI

Related searches

Same AI level