Senior Machine Learning Engineer (Large Systems)
AI in this role
About Graphcore
At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence.
Job Summary
As a Senior Machine Learning Engineer in the Applied AI team at Graphcore, you will contribute to advancing AI technology by developing and optimising AI models tailored to our specialised hardware. You will work on large scale systems where performance is critical to the success of our projects. Working closely with the Software development and Research teams, you will play a critical role in identifying opportunities to innovate and differentiate Graphcore’s technology. We seek engineers with strong technical skills and an understanding of AI model implementation at scale, eager to make a tangible impact in this rapidly evolving field.
The Team
The Applied AI team’s role is to be proxies for our customers, we need to understand the latest AI models, applications, and software to ensure that Graphcore’s technology works seamlessly with the AI ecosystem and at scale. We build reference applications, contribute to key software libraries e.g. optimising kernels for efficiency on our hardware, and collaborate with the Research team to develop and publish novel ideas in domains such as efficient compute, model scaling and distributed training and inference of AI models for multiple modalities and applications.
If you're excited about advancing the next generation of AI models on cutting-edge hardware, we’d love to hear from you!
Responsibilities and Duties
- Implement latest machine learning models and optimise them for performance and accuracy, scaling to 1000s of accelerators.
- Test and evaluate new internal software releases, provide feedback to software engineering teams, make necessary code fixes, and conduct code reviews.
- Benchmark models and key ML techniques to identify performance bottlenecks and improve model efficiency.
- Design and conduct experiments on novel AI methods, implement them and evaluate results.
- Collaborate with Research, Software, and Product teams to define, build, and test Graphcore’s next generation of AI hardware.
- Engage with AI community and keep in touch with the latest developments in AI.
Candidate Profile
Essential:
- Bachelor/Master's/PhD or equivalent experience in Machine Learning, Computer Science, Maths, Data Science, or related field.
- Proficiency in deep learning frameworks like PyTorch/JAX.
- Strong Python or C++ software development skills
- Expertise in deep learning from model training to optimisation and evaluation.
- Capable of designing, executing and reporting from ML experiments.
- Developed deep understanding of performance bottlenecks and how to overcome them.
- Ability to move quickly in a dynamic environment
- Enjoy cross-functional work collaborating with other teams.
- Strong communicator - able to explain complex technical concepts to different audiences.
Desirable:
- Experience in one or more of:
- MLOps for Kubernetes-based clusters
- Building production systems with large language models
- Efficient computing based on low-precision arithmetic.
- Experience writing C++/Triton/CUDA kernels for performance optimisation of ML models.
- Experience in distributed training or inference of ML models across 64+ accelerators.
- Familiarity with HPC systems and networking including Infiniband, NVLink, RoCE technologies.
- Have contributed to open-source projects or published research papers in relevant fields.
- Knowledge of cloud computing platforms.
- Keen to present, publish and deliver talks in the AI community.
Benefits
In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications
How we rate this
Senior Machine Learning Engineer (Large Systems) at Graphcore rates 99 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you monitor a model once it's live, and how do you know it needs retraining?
- Walk me through how you've used PyTorch in your day-to-day work.
- What are the limits of Jax that you've run into, and how did you work around them?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: ML Ops, PyTorch, and Jax. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new machine learning engineer jobs (Builds AI ●●●●) by email
One email a week with the new machine learning engineer jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Data roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step