We're the team behind Latent Diffusion, Stable Diffusion, and FLUX — foundational technologies that changed how the world creates images and video. Our models power the tools used by millions of creators, developers, and businesses worldwide, and FLUX is among the most advanced generative systems in the world.
Headquartered in Freiburg, Germany with a growing presence in San Francisco, we're scaling fast while staying true to what makes us different: research excellence, open science, and building technology that expands human creativity.
Vision-language models are becoming foundational to how people interact with generative AI — but most VLM research happens in isolation from the generation stack. At Black Forest Labs, we're integrating VLMs directly into FLUX in ways that make our models more powerful, more controllable, and more aligned with what creators actually want.
This role is about pioneering that integration. You won't be applying off-the-shelf VLMs — you'll develop novel approaches, innovate on architectures, and answer questions that haven't been solved yet: how vision and language representations inform each other, how multimodal understanding improves generation quality, and how to make these capabilities deployable at scale without compromising what makes FLUX exceptional.
This is a Staff / Senior IC role. We're looking for someone who has pretrained or significantly advanced a VLM, not just fine-tuned one.
We’re a distributed team with real offices that people actually use. Depending on your role, you’ll either join us in Freiburg or SF at least 2 days a week (or one full week every other week), or work remotely with a monthly in-person week to stay connected. We’ll cover reasonable travel costs to make this possible. We think in-person time matters, and we’ve structured things to make it accessible to all. We’ll discuss what this will look like for the role during our interview process.
Everything we do is grounded in four values:
If this sounds like work you’d enjoy, we’d love to hear from you.
Base Annual Salary:
EU €130,000-€340,000 + Equity
Find more English Speaking Jobs in Germany on Arbeitnow
You'll thrive on the deep intellectual challenge of pioneering novel VLM architectures and integrating them with diffusion models.
Read the INTJ career guide →Your natural curiosity and love for theoretical exploration make you ideal for innovating on multimodal architectures and evaluating emerging research.
Read the INTP career guide →Your drive to lead and execute makes you a strong fit for driving the development of state-of-the-art VLMs from concept to deployment.
Read the ENTJ career guide →