
Inception
Diffusion-based LLMs
Inception is the team behind Mercury, the first productionized family of diffusion-based large language models. Instead of producing one token at a time like every other LLM you've used, Mercury generates many tokens in parallel — which makes it several times faster, at a fraction of the cost of comparable autoregressive models.
They are a small team making an audacious bet. Their founders pioneered diffusion modeling, Flash Attention, and Direct Preference Optimization, among others. The team blends researchers and engineers from Stanford, Cornell, UCLA, Google DeepMind, Meta AI, Microsoft, and OpenAI. If you want to work close to the frontier of a new generative paradigm, this is the place to be.