Sword HealthRemote · Remote - Portugal€50k-€72k
AmazonPosted 1mo ago
L4
ML Software Engineer, Data Plane
ML Software Engineer, Data Plane at Amazon scores 98 out of 100 on AI centrality, which makes it a Level 4 role on this board.
IL, Tel Avivfull-time
AI in this role
vllmpytorchjax
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.
Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Basic qualifications
- Bachelor's degree or equivalent
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Knowledge of computer architecture, operating systems, and parallel computing
- Strong proficiency in C/C++
- Strong Linux systems knowledge
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators
- Proven track record of owning and delivering complex software features end-to-end
Preferred qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
vLLMPyTorchJax
Questions you could be asked
- What's a project where you used vLLM hands-on?
- Walk me through how you've used PyTorch in your day-to-day work.
- What are the limits of Jax that you've run into, and how did you work around them?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: vLLM, PyTorch, and Jax. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Similar roles
Software Engineering roles rated Level 4 at other companies.




