Text Annotation Engineer
Munich, Germany · Zürich, Switzerland
Posted on Aug 8, 2026
The next ten years of AI will not be won in software. They will be won in the physical world. In factories, hospitals, kitchens, fields, and homes. The companies that own that data will own the century.
microagi is building it. We are the data layer for physical AI.
Nobody owns this stage yet. It is yours to build. Every clip we collect needs language attached to it. What the person is doing, what the task is, where one step ends and the next begins. You will build that with vision-language models where they are good enough, and with people where they are not.
What You Will Do
- Build the text annotation stage from nothing. Captions, task descriptions, step segmentation.
- Use vision-language models for the first pass and measure how often they are right.
- Design the human-in-the-loop review that corrects what the models get wrong.
- Write the guidelines annotators follow, then rewrite them when they turn out to be ambiguous.
- Build quality metrics for language output, which is harder to score than a bounding box.
- Scale the stage from a few hours of footage to thousands.
Requirements
- Strong Python and hands-on work with vision-language models or LLMs.
- Experience building annotation or labeling workflows, including the human side of them.
- You can define what counts as a correct caption and then measure it.
- Comfortable owning a stage that does not exist yet.
- High agency - you don't wait to be told what to do.
- Fluent in English.
Nice to Have
- Video understanding or temporal action segmentation.
- Fine-tuning or evaluating VLMs.
- Experience managing external annotation teams.
#LI-DNI