Research Scientist (Reinforcement Learning)
Munich, Germany · Zürich, Switzerland
Posted on Aug 18, 2026
The next ten years of AI will not be won in software. They will be won in the physical world. In factories, hospitals, kitchens, fields, and homes. The companies that own that data will own the century.
microagi is building it. We are the data layer for physical AI.
A policy that succeeds forty percent of the time is a demo. A policy that succeeds 99.99 percent of the time is a product. Closing that gap is the whole job.
You will post-train robot policies with online and offline reinforcement learning, and you will own the number that says whether it worked.
What You Will Do
- Post-train robot policies with online and offline reinforcement learning and related methods.
- Take policies from a demo-grade success rate to one a customer would accept, and know exactly why every point of it moved.
- Build and tune simulation environments, then close the gap to the real robot.
- Build the reward, evaluation, and logging infrastructure the experiments depend on.
- Run experiments daily. Keep what holds up and throw away the rest.
- Take a policy from simulation onto hardware and make it survive contact with the physical world.
Requirements
- Deep reinforcement learning experience across both online and offline methods. You have trained policies that worked, and you know why the others did not.
- You have moved a real success rate a long way, and you can account for every step of how.
- Experience with robot simulation and sim-to-real transfer. Hands-on with Isaac Sim, MuJoCo, ManiSkill, or equivalent.
- Strong Python and PyTorch or JAX.
- Publications at strong venues are welcome but not required. Results are.
- You can hold a direction for months and still kill it when the evidence says to.
- High agency - you don't wait to be told what to do.
- Fluent in English.
Nice to Have
- Experience training vision-language-action models or diffusion policies.
- Experience with very large training runs, including the GPU orchestration behind them.
- Open-source work that other researchers build on.
#LI-DNI