Meet Iris-3B: a pixel-space T2I model and general vision learner that generates every pixel directly. No VAE.
Pre-trained from scratch, 3B, open weights.
We also explore its generative prior on detail-critical vision tasks: depth estimation and image restoration.
Fully open: