Research Scientist (Omni-Model Pretraining)

microagi
microagi

Zürich, Switzerland · London, UK

Posted on Aug 18, 2026

The next ten years of AI will not be won in software. They will be won in the physical world. In factories, hospitals, kitchens, fields, and homes. The companies that own that data will own the century.

microagi is building it. We are the data layer for physical AI.

In eight months we will be sitting on a million hours of expert-annotated human data. Nobody annotates that by hand. The models that describe what a person is doing, where one step ends and the next begins, and what the world looked like while they did it have to be trained by us, on our own data.

That is a research problem, not a compute problem. This role owns it. Omni-model pretraining across egocentric video, depth, IMU, audio, and language, and the world-action models we build on top of it.

You will be a sparring partner for people who have been doing this a long time. What to train, which architecture, what to throw away. We want someone with real opinions about what great research takes, and the record to back them.

What You Will Do

  • Pretrain omni and video models at scale on egocentric capture. Video, depth, IMU, audio, and language together.
  • Build the world-action and omni models that learn how the physical world behaves from watching people work in it.
  • Train the vision-language models that annotate our data. Captions, task descriptions, step segmentation, and the long tail nobody has a label schema for yet.
  • Design the human-in-the-loop review that corrects what the models get wrong, and shrink it every quarter.
  • Build quality metrics for language and multimodal output, which is far harder to score than a bounding box.
  • Own data preparation at our scale. Curation, filtering, mixing, and dedup decide the result more often than architecture does.
  • Build the evaluation sets that tell you whether a result is real or a measurement artifact.
  • Turn results into things that ship. Better annotations, better datasets, better models for customers.

Requirements

  • You have trained large models yourself, end to end, and you know what breaks at scale.
  • Depth in omni-model or video pretraining, video generation models, or vision-language models. It does not have to be robotics.
  • Serious data preparation experience. You have made the curation and mixing calls on a large training corpus.
  • Strong PyTorch or JAX. You write the training code yourself.
  • Large-scale distributed training. You are comfortable with the infrastructure, not only the model.
  • You can hold your own on model architecture with people who have shipped frontier models.
  • Publications at strong venues, models in production, or independent work of the same calibre. We do not care where you did it.
  • You can hold a direction for months and still kill it when the evidence says to.
  • High agency - you don't wait to be told what to do.
  • Fluent in English.

Nice to Have

  • Experience with egocentric datasets such as Ego4D, Ego-Exo4D, or EPIC-Kitchens.
  • Experience with video understanding or temporal action segmentation.
  • Familiarity with hand, body, and object understanding, including SMPL, SMPL-X, or MANO.
  • Familiarity with depth estimation, 6-DoF object pose, or 3D reconstruction.
  • Experience building annotation pipelines and measuring their quality.
  • Open-source work that other researchers build on.

#LI-DNI