Just skimmed the paper (haven't seen this before). Might need to read more carefully, but at first glance I don't understand their intuition why the "Markovian property limits the model’s ability to fully utilize the generation trajectory".
When it comes to multi-resolution training (e.g. matryoshka training), there are precedents that don't require this AR formulation.
They cite MAR (https://arxiv.org/pdf/2406.11838), which I think is a much clearer articulation of "autoregressive diffusion". MAR uses an autoregressive base (like an LLM) and staples on a small MLP on-top that's trained as a diffusion head. That makes more sense to me, since you can leverage the "knowledge prior" from an LLM and have it generate images. That's probably how nano-bannana and GPT-Image broadly work.
Zooming out, it's not clear to me from any of the work WHY autoregressive diffusion on it's own is better than regular diffusion. Most autoregressive diffusion models use speculative decoding, because pure autoregressive diffusion is too slow to run at inference time.
The one ~magical~ thing about auto-regressive diffusion (IMO) has nothing to do with autoregressive vs. fully-bidirectional diffusion. But simply, the face that you can do it on-top of a LLM. That means the LLM can look at it's generation, assess it's quality, think on what's broken, and then call itself to edit the image and fix it. With methods like RLVR, that means you can essentially guarantee the correctness of your image (along the axes you've RLVR-ed).
In my opinion its a great idea because you get reuse hidden states which allows you to do something akin to thinking/[online learning]. If you only evolve use input space outputs you're dealing with more decoding pressure.
The HN poster of this owns an extremely anti-social page btw https://boomerdeathwatch.com/ . Before he claims that "they deserved it", note how the page itself does nothing but make them look like a weirdo lol
Here's another meta-lesson: "woke" turns out to be broadly popular and is not actually going away.
Yes, many people *genuinely* give a fuck that tech leaders are promoting white supremacy and ethnic cleansing, and there will be actual consequences for this.
I really want to like ai music but none of them seem to be trained on reward models that reward "ambiance" or any sort of interesting sound design (which this paper obviously doesn't even concern), which might be reflective of the people training them not having niche music tastes
I don't share my friend's music taste and record stores for sure don't have what I like (its usually small soundcloud accounts). Also the things I like about a song are not genre bound, its usually very subtle things that are unsearchable, hence even of the songs I like I usually only like ~30% of the song itself. With a personally tuned reward model I can strictly focus on amplifying the subtleties instead of searching by means of exhausting indirection
Even normal fridges are a memetic virus, fridges clearly live in a symbiotic relationship with humans. They make our lives better and in return we create more fridges. If they weren't useful to us, we'd stop making more fridges. Any medieval human who sees a fridge would instantly recognize why it's useful and valuable.
OP is right, anything popular is virus-like. Richard Dawkins invented the word meme for this all the way back in 1976.
reply