AlmetraPosted 9mo ago
VLM Research Engineer (m/f/d) at Almetra scores 95 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI in this role
We’re looking for a Research Engineer to push the limits of vision-language models for real-world video understanding. You’ll work on applied, state-of-the-art multimodal models and turn them into production pipelines used by customers.
Your role
Design and adapt vision-language and video models for scene understanding, temporal reasoning and activity / action recognition
Build and maintain large-scale training and evaluation pipelines on GPU clusters
Curate and augment video-text and action datasets, including synthetic labels and retrieval-based augmentation
Develop robust benchmarks for video QA, instruction following and temporal understanding, and use them to drive iterative model improvements
Cut and refactor model architectures for efficiency and deployability (compression, pruning, distillation)
Deliver production-ready inference pipelines to product and customer teams, working closely with CV, platform and robotics engineers
You bring
Completed PhD (or equivalent research track record) in computer vision, machine learning, robotics or a related field
Strong background in video-centric deep learning: scene understanding, temporal / activity / action recognition, or video generation
Experience training and adapting large vision or VLM models (e.g. InternVL, Qwen-VL, DeepSeek-VL, similar stacks)
Proven work with multi-GPU training (PyTorch, distributed, mixed precision) and large-scale datasets
Solid engineering habits: clean Python, reproducible experiments, reliable data and training pipelines
Track record of moving research into usable systems (demos, internal tools, or productised features) in fast-moving teams
Nice to have
Publications at top-tier venues (CVPR, ICCV, ECCV, NeurIPS, ICLR, etc.) on video, multimodal learning or scene understanding
Experience with 3D/4D scene representations, action generation or embodied / sense-plan-act style projects
Inference optimisation: quantisation, TensorRT, model distillation, or deployment on constrained hardware
Prior experience in a startup or applied research lab environment
What we offer
Employee Share Options Program for all permanent employees*
An increasing benefits list: currently includes Urban Sports club and quarterly team retreats.
Be on the forefront in defining what artificial intelligence means in manufacturing
Gain hands-on experience in working in an AI-first software company
Supportive and inclusive culture that values diversity and promotes the advancement of underrepresented groups within the company
Collaborate with a diverse (currently more than 10 nationalities) and talented team, working on cutting-edge projects with real-world impact
Network with professionals and leaders in the field, opening doors to potential future career opportunities
We have a very flat hierarchy, open 360° feedback, and flexible working hours
Ethics⚖: We are committed to developing ethical AI software
Don't meet all the requirements?
Almetra is committed to creating a workplace that is diverse, fair, and inclusive. We encourage candidates from all backgrounds, even if they do not meet every qualification, to submit their application. We firmly believe that having a team with diverse perspectives only strengthens our company and drives innovation. Our commitment also extends to providing an accessible environment for everyone, including those with disabilities. Please let us know if you require any accommodations during the application process or while working with us, and we will do our best to support you.
*Only full-time, permanent roles are eligible for stock options. Part-time roles, contract roles, work-student, internships and freelance roles are not eligible for stock options.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Walk me through a computer vision problem you solved, from raw data to a deployed model.
- Walk me through how you've used Deepseek in your day-to-day work.
- What are the limits of PyTorch that you've run into, and how did you work around them?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: Computer Vision, Deepseek, and PyTorch. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Similar roles
Research roles rated AI Level 4 at other companies.
NVIDIAUS, CA, Santa Clara$38-$94/hr
MercorRemote · San Francisco$5000k
WaymoRemote · Mountain View, CA, USA; San Francisco, CA, USA; New York, NY, USA$213k-$263k
Anduril IndustriesBroomfield, Colorado, United States; Fort Collins, Colorado, United States$190k-$252k






