Training a 210M text-to-image DiT on one GPU
The author trained a 210M-parameter DiT on one RTX PRO 6000 (3.5 days, ~4.2M 256² images) and reports three findings: learned null attention slots dominate cross-attention, flow-matching loss reflects training health not quality, and a training-time timestep shift notably improves sampling quality.