Overview
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead.
- We propose Embedding Conditioned Generation, which conditions a DiT generator on the clean embeddings predicted by a NEPA model from the condition and the noisy image, recomputed at every denoising step.
- We train the NEPA model with Multi-Embedding Prediction, which extends NEPA to predict the embeddings of the whole clean image at once.
- On class-conditional ImageNet 256Γ256, we study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models; combined with REPA, NEPA-DiT-XL reaches an FID of 1.32.
Method
Our method has two stages. First, we train a NEPA model with Multi-Embedding Prediction to predict clean image embeddings from condition tokens and noisy image embeddings. Second, we train a DiT generator with Embedding Conditioned Generation, where the frozen NEPA model produces predictive embeddings at each denoising step and the DiT uses them as its condition.
Multi-Embedding Prediction. Multi-Embedding Prediction (MEP) extends NEPA to predict the next embeddings simultaneously from the same context.
NEPA is the special case . Like multi-token prediction, MEP predicts several items at once; it differs in predicting in embedding space rather than token space.
NEPA for Generation. In this sequence, the clean embeddings are the next embeddings after the condition and the noisy image, and we train a NEPA model to predict them with MEP, taking . The sequence is ordered by generation state rather than by patch position, so the embedding that follows noisy patch is clean patch . MEP predicts all embeddings of the clean image in one forward pass, each from the full context of the condition and the noisy image.
Attention mask. The condition tokens and form a causal prefix that never attends to image tokens.
Embedding prediction loss. As illustrated in Figure 5, the positives lie on the diagonal, pairing each with , while the other patches of the same image serve as negatives:
Embedding Conditioned Generation. At each denoising step, the current noisy input is first fed into the NEPA model, which outputs predictive embeddings .
The generator receives the condition only through and has no separate class embedding; the timestep enters through adaLN modulation, as in DiT.
Condition caching. Because the condition prefix never attends to image tokens, its key-value cache in the NEPA model is computed once before the denoising loop and reused at every step; only the image tokens are re-encoded as changes.
Classifier-free guidance. During training, we randomly replace the condition with a learned null token. At inference, the unconditional branch feeds the null token to the NEPA model, which yields , and guidance is applied to the velocities predicted by the generator:
Ablation Studies
Following previous work (Ma et al., 2024; Wang et al., 2026b), we run ablations at the B scale with NEPA-B and patch size 2, one factor at a time (defaults in gray); Table 6 then selects patch size 4.
Loss. InfoNCE gives the best FID, 31.36 against 35.23 for MSE and 37.82 for cosine similarity (Table 1). ImageNet-1K accuracy of the fine-tuned NEPA model follows the opposite order.
| Loss | FIDβ | Acc. (%)β |
|---|---|---|
| Cos. sim. | 37.82 | 83.1 |
| MSE | 35.23 | 82.6 |
| InfoNCE | 31.36 | 81.7 |
| Detach | FIDβ |
|---|---|
| w/ | n.c. |
| w/o | 31.36 |
Stop-gradient. Detaching the targets from the gradient prevents training from converging (n.c. in Table 2), so gradients flow through the targets.
Target normalization. Normalizing the target embeddings raises FID from 31.36 to 34.76 (Table 3), so we use raw targets.
Timestep input. Feeding the diffusion timestep to the NEPA model changes FID from 31.36 to 31.40 (Table 4), presumably because the noise level is already evident from the noisy patches; we therefore keep the NEPA model timestep-free, while the generator still receives the timestep through adaLN.
| Target | FIDβ |
|---|---|
| Norm. | 34.76 |
| Raw | 31.36 |
| Timestep | FIDβ |
|---|---|
| w/o | 31.36 |
| w/ | 31.40 |
Vocabulary. The InfoNCE loss scores each prediction against a bank of candidate targets. We compare banks built from the patches of the same image (instance), of all images in the per-device batch (batch), and of all images across devices (full). The instance, batch, and full banks give 31.36, 32.13, and 31.21 (Table 5). We use the instance bank, which requires no communication across devices; its negatives are the other patches of the same image, as in Equation 5.
Patch size. The patch size of the NEPA model sets the number of predicted embeddings (Table 6). Patch size 4 gives 31.00 against 31.36 for patch size 2 with a quarter of the tokens, which makes the query at every denoising step correspondingly cheaper, whereas patch size 8 is too coarse (33.20). We therefore use patch size 4 in all remaining experiments.
| Bank | FIDβ |
|---|---|
| Instance | 31.36 |
| Batch | 32.13 |
| Full | 31.21 |
| Patch | Tokens | FIDβ |
|---|---|---|
| 2 | 256 | 31.36 |
| 4 | 64 | 31.00 |
| 8 | 16 | 33.20 |
Condition of the Generator
At the same small scale, Table 7 keeps the generator and its training recipe fixed and varies only what it is conditioned on. The class embedding gives 36.39, and adding the patch embeddings of to the class token gives 37.21. The last three rows give the generator the same extra network as ours, of the NEPA-XL architecture, so that their parameters and FLOPs per step match those of our model; what differs is how that network is trained. Trained end to end with the generator as a single network, it gives 29.52. Pretrained as a flow-matching model for 240 epochs, the same budget as MEP, and then frozen, its features give 30.86. Trained with MEP and frozen, its predicted embeddings give 25.04.
| Condition of the generator | Extra network | Training | Objective | FIDβ |
|---|---|---|---|---|
| Class embedding | β | β | β | 36.39 |
| Class token + noisy patches | β | β | β | 37.21 |
| Single network, end to end | NEPA architecture | joint | flow matching | 29.52 |
| Pretrained flow model | NEPA architecture | 240 epochs, frozen | flow matching | 30.86 |
| NEPA model | NEPA | 240 epochs, frozen | MEP | 25.04 |
Refresh rate. Table 14 queries the NEPA model less often than every denoising step, reusing the last prediction in between. Querying it every 8 steps raises FID from 1.57 to 1.82, and querying it only once, at the first step, gives 242.63.
| NEPA query | FIDβ |
|---|---|
| Every step | 1.57 |
| Every 8 steps | 1.82 |
| Once, at the first step | 242.63 |
Scaling Behavior
Beyond the small-scale setting, Table 8 scales both models, training all nine combinations of NEPA-B/L/XL and DiT-B/L/XL, each for 400K iterations and evaluated with 64 ODE steps without guidance. FID decreases monotonically along both axes. Scaling the generator from DiT-B to DiT-XL improves FID by about 14 points for every NEPA model, and scaling the NEPA model from NEPA-B to NEPA-XL improves FID by about 5 points for every generator.
| Generator (FIDβ) | ||||
|---|---|---|---|---|
| NEPA | Acc.β | DiT-B | DiT-L | DiT-XL |
| NEPA-B | 81.3 | 31.00 | 17.69 | 16.54 |
| NEPA-L | 83.4 | 27.46 | 14.35 | 13.74 |
| NEPA-XL | 85.2 | 25.04 | 12.86 | 11.27 |
Training progress. Figure 6 follows NEPA-DiT-B/L/XL over training; all three use the same frozen NEPA-XL model. Each point is FID-50K with 64 ODE steps and no guidance, evaluated every 100K iterations. Larger generators are better at every checkpoint, and the gap opens early: at 200K iterations, NEPA-DiT-L (17.55) and NEPA-DiT-XL (15.23) are already well ahead of NEPA-DiT-B at 400K (25.04).
Comparison with Previous Methods
Our final model combines Embedding Conditioned Generation, which sets the condition of the generator, with REPA (Yu et al., 2025), which aligns its features.
Among the methods of Table 9 in the SD-VAE latent space, NEPA-DiT-XL + REPA reaches an FID of 1.32 with the 250-step SDE sampler and 1.57 with the 96-step ODE sampler. Most of its training compute goes into the NEPA model, trained once: with it included, NEPA-DiT-XL, 1.39B parameters in all, is trained with 3.1Γ1020 FLOPs, about a third of what SiT-XL/2 with REPA uses over its 800 epochs to reach 1.42.
| Method | Tokenizer | Epochs | FIDβ | sFIDβ | ISβ | Pre.β | Rec.β |
|---|---|---|---|---|---|---|---|
| Other tokenizers | |||||||
| REPA + EQ-VAE (Kouzelis et al., 2025a) | EQ-VAE | 200 | 1.70 | 5.13 | 283.0 | 0.79 | 0.62 |
| LightningDiT-XL/1 (Yao et al., 2025) | VA-VAE | 800 | 1.35 | 4.15 | 295.3 | 0.79 | 0.65 |
| LightningDiT + IG (Zhou et al., 2026) | VA-VAE | 680 | 1.19 | 4.11 | 269.0 | 0.79 | 0.66 |
| DiT-XL + CMuon (Chen et al., 2026) | VA-VAE | 200 | 1.18 | β | β | β | β |
| REPA-E (Leng et al., 2025) | E2E-VAE | 800 | 1.12 | 4.09 | 302.9 | 0.79 | 0.66 |
| Send-VAE + REPA (Page et al., 2026) | Send-VAE | 800 | 1.21 | 4.10 | 315.1 | 0.79 | 0.66 |
| SFD-XLβ‘ (Pan et al., 2026b) | SemVAE | 800 | 1.06 | 3.89 | 267.0 | 0.78 | 0.67 |
| RAE DiTDH-XLβ‘ (Zheng et al., 2026) | RAE | 800 | 1.13 | β | 262.6 | 0.78 | 0.67 |
| MixFlow + RAE (Li et al., 2026) | RAE | 800+200 | 1.10 | 4.40 | 259.7 | 0.78 | 0.67 |
| RAEv2 (Singh et al., 2026) | RAE | 80 | 1.06 | β | 255.3 | β | β |
| SD-VAE | |||||||
| DiT-XL/2 (Peebles & Xie, 2023) | SD-VAE | 1400 | 2.27 | 4.60 | 278.2 | 0.83 | 0.57 |
| SiT-XL/2 (Ma et al., 2024) | SD-VAE | 1400 | 2.06 | 4.49 | 277.5 | 0.83 | 0.59 |
| REPA (Yu et al., 2025) | SD-VAE | 800 | 1.42 | 4.70 | 305.7 | 0.80 | 0.65 |
| TREAD (Krause et al., 2025) | SD-VAE | 740 | 1.69 | 4.73 | 292.7 | 0.81 | 0.63 |
| DDT-XL/2 (Wang et al., 2026b) | SD-VAE | 400 | 1.26 | β | 310.6 | 0.79 | 0.65 |
| REG (Wu et al., 2025) | SD-VAE | 800 | 1.36 | 4.25 | 299.4 | 0.77 | 0.66 |
| U-REPA (Tian et al., 2026) | SD-VAE | 400 | 1.41 | β | β | β | β |
| SRA (Jiang et al., 2026) | SD-VAE | 800 | 1.58 | 4.65 | 311.4 | 0.80 | 0.63 |
| LSEP (Yun et al., 2025) | SD-VAE | 800 | 1.46 | 4.94 | 296.8 | 0.80 | 0.64 |
| SPRINT + REPA (Park et al., 2026) | SD-VAE | 400 | 1.59 | β | β | 0.80 | 0.64 |
| SiT-XL/2 + IG (Zhou et al., 2026) | SD-VAE | 800 | 1.46 | 4.79 | 265.7 | 0.80 | 0.64 |
| SiT-XL/2 + Sparse Guid. (Krause et al., 2026) | SD-VAE | 400 | 1.58 | 4.45 | 249.7 | 0.80 | 0.63 |
| SRA 2 (Wang et al., 2026a) | SD-VAE | 800 | 1.52 | 4.63 | 316.2 | 0.82 | 0.62 |
| UDT+-XL/2 + REPA (Yun et al., 2026) | SD-VAE | 320 | 1.38 | 4.38 | 306.3 | 0.79 | 0.66 |
| RecFM-XL (Huang et al., 2026) | SD-VAE | 160 | 2.49 | β | β | β | β |
| NEPA-DiT-XL + REPA (ours), ODE-96 | SD-VAE | 240+80 | 1.57 | 4.82 | 298.0 | 0.79 | 0.63 |
| NEPA-DiT-XL + REPA (ours), SDE-250 | SD-VAE | 240+80 | 1.32 | 4.32 | 311.0 | 0.79 | 0.68 |
Guidance. Restricting guidance to an interval (KynkÀÀnniemi et al., 2024) admits larger scales: scale 3.6 on gives 1.57 and 1.32, the best with both samplers, selected by FID-50K.
Query rate. Querying at every step also sets the sampling cost of the NEPA model, 92 GFLOPs per step next to 309 for the generator; with 96 ODE steps, one image takes 62 TFLOPs in all, below the 91 TFLOPs of SiT-XL/2 with REPA with 250 SDE steps.
| Scale | Interval | ODE-96 | SDE-250 |
|---|---|---|---|
| 1.0 | [0, 1] | 8.77 | 7.59 |
| 1.4 | [0, 1] | 2.58 | 2.32 |
| 2.4 | [0.3, 1] | 1.72 | 1.46 |
| 3.6 | [0.4, 1] | 1.57 | 1.32 |
| Params | GFLOPs / step | ||||
|---|---|---|---|---|---|
| Model | Generator | NEPA | Generator | NEPA | TFLOPs / image |
| NEPA-DiT-B | 135M | 711M | 59 | 92 | β |
| NEPA-DiT-L | 458M | 711M | 210 | 92 | β |
| NEPA-DiT-XL, ODE-96 | 683M | 711M | 309 | 92 | 62 |
| NEPA-DiT-XL, SDE-250 | 683M | 711M | 309 | 92 | 160 |
| SiT-XL/2 + REPA, SDE-250 | 675M | β | 229 | β | 91 |
Additional Qualitative Results
Figures 8β10 show more samples of NEPA-DiT-XL on ImageNet 256Γ256, selected from 16 samples per class.
























Conclusion and Future Work
We have shown that predictive embeddings can serve as the condition of an image generator. In Embedding Conditioned Generation, a DiT generator is conditioned at every denoising step on the clean-image embeddings that a NEPA model, trained with Multi-Embedding Prediction, predicts from the condition and the current noisy image. On ImageNet 256Γ256, combined with REPA, NEPA-DiT-XL reaches an FID of 1.32 with the 250-step SDE sampler and 1.57 with the 96-step ODE sampler, trained with about a third of REPA's compute even though it trains and runs two networks.
Text conditions. Text-to-image models either encode the prompt once with a separate text encoder, which cannot adapt to the evolving noisy image and often fixes the length of the condition, or process the prompt together with the image in a single unified model, where the condition and the image share one backbone. A NEPA model could instead read the whole prompt before the noisy image and predict the clean image embeddings from both, as it does for class labels in this paper.
Unified models. Since the NEPA model is an autoregressive Transformer, a single NEPA model could in principle serve both as a language model with an output head and as the conditioning model of a DiT generator. Studying text conditions, other latent spaces, and higher resolutions are directions for future work.