JiT-DDT cuts text-to-image training time 3.6×
Researchers propose JiT-DDT, an encoder-decoder architecture that trains text-to-image models more efficiently than LDMs. Compared to a Linum v2 baseline, it claims 3.6× fewer GPU-hours while producing higher-resolution images; code and weights are released under Apache 2.0.