Research Scientist (Computer Vision)
Software Engineering
Zürich, Switzerland · London, UK
Posted on Jul 31, 2026
The next ten years of AI will not be won in software. They will be won in the physical world. In factories, hospitals, kitchens, fields, and homes. The companies that own that data will own the century.
MicroAGI is building it. We are the data layer for embodied AGI.
You will own and advance the perception and annotation stack that turns raw egocentric capture (RGB, RGB-D, IMU, stereo, audio) into high-fidelity training data. The role spans research and production: improving core computer-vision modules, benchmarking them rigorously, and shipping them into the pipeline.
What You Will Do
- Build and improve hand, body, and object-tracking pipelines on egocentric and multi-view footage.
- Advance 3D pose and shape estimation, SLAM, depth estimation, segmentation, and adjacent perception modules.
- Fuse visual and IMU signals under occlusion, motion blur, and real-world capture conditions.
- Turn papers, prototypes, and internal experiments into reliable production improvements.
- Build evaluation sets and benchmark annotations against academic datasets and customer requirements.
- Work across research, engineering, and delivery to make the annotation stack accurate, robust, and scalable.
Requirements
- Strong foundation in computer vision, machine learning, and deep learning.
- Hands-on experience with tracking, pose estimation, SLAM, depth, segmentation, or visual-inertial fusion.
- Strong mathematical foundation and comfort reading research papers.
- Publications at strong venues or production work that demonstrates equivalent depth.
- Careful engineering judgment around production pipelines.
- High agency - you don't wait to be told what to do.
- Fluent in English.
Nice to Have
- Direct experience with SMPL, SMPL-X, MANO, or adjacent body and hand models.
- IMU-based motion capture or visual-inertial fusion.
- Experience with annotation pipelines, quality assurance, or dataset evaluation.
- Robotics experience in imitation learning, teleoperation, or manipulation.
#LI-DNI