A team of artificial intelligence researchers—comprising Jiatao Gu, Tianrong Chen, Ying Shen, David Berthelot, Shuangfei Zhai, and Josh Susskind—has introduced a novel generative framework known as Normalizing Trajectory Models (NTM). The newly proposed architecture aims to address a fundamental limitation in modern diffusion-based models: the degradation of underlying mathematical assumptions when the generation process is compressed from numerous small steps into a handful of coarse transitions. Diffusion models have dominated the landscape of generative artificial intelligence, particularly in text-to-image synthesis, by decomposing the sampling procedure into a sequence of many small Gaussian denoising steps. While this iterative refinement yields exceptionally high-fidelity outputs, it comes at a significant computational cost. Generating a single sample often requires dozens or even hundreds of sequential forward passes through a neural network, making real-time applications challenging. Read Also: Unintended Consequences: New Study Reveals How Value Induction Makes Large Language Models More Sycophantic and Addictive New Research Breakthrough Narrows Performance Gap in Semi-Supervised Federated Learning for Automatic Speech Recognition To accelerate this process, existing few-step generation methods typically rely on various post-processing techniques such as distillation, consistency training, or adversarial objectives. Although these approaches successfully reduce the required sampling steps, they invariably force practitioners to sacrifice the likelihood framework. In doing so, these models lose the rigorous probabilistic tracking and exact density estimation that make likelihood-based models mathematically appealing and stable to train. The introduction of Normalizing Trajectory Models seeks to bridge this longstanding divide. NTM models each reverse step of the generation process as an expressive conditional normalizing flow, thereby maintaining exact likelihood training. By marrying the strengths of normalizing flows with the trajectory-based generation typical of modern diffusion systems, the authors have developed a framework that preserves rigorous probabilistic guarantees while drastically cutting down the number of steps required to produce high-quality outputs. From an architectural standpoint, NTM is designed to balance local flexibility with global trajectory coordination. The network combines shallow invertible blocks within each individual generation step alongside a deep parallel predictor that spans across the entire trajectory. This dual structure forms a cohesive, end-to-end network that can either be trained entirely from scratch or conveniently initialized from pre-trained flow-matching models, providing flexibility for researchers and engineers working with existing infrastructures. A key capability enabled by the exact trajectory likelihood of NTM is self-distillation. Through this mechanism, a lightweight denoiser is trained directly on the score function induced by the model itself. This self-guided distillation process ultimately produces high-quality samples in just four sampling steps, circumventing the need for complex external teacher models or adversarial training loops. When evaluated on standard text-to-image benchmarks, NTM demonstrates competitive performance. The model matches or outperforms strong existing image generation baselines while requiring only four sampling steps. Crucially, it achieves this high operational efficiency while uniquely retaining the property of exact likelihood over the generative trajectory, a dual achievement that sets it apart from conventional accelerated diffusion or consistency models. Related readings and updates The broader landscape of generative modeling continues to see active exploration into the mechanics of multi-step trajectories and alternative probabilistic frameworks. Among related developments, discrete flow matching has emerged as a prominent technique for generating text by iteratively transforming noise tokens into coherent language. Much like its continuous counterpart in image generation, discrete flow matching relies on transforming distributions through forward passes, sometimes numbering in the hundreds. To address these computational bottlenecks, researchers frequently employ distillation techniques, using a multi-step trajectory as a teacher to train a student model capable of reproducing the generation process in significantly fewer steps. When student models occasionally underperform in these setups, conventional intuition often attributes the limitation to insufficient model capacity. However, ongoing research challenges this assumption, suggesting that the bottleneck frequently lies within the training trajectory itself rather than the capacity of the student network, as each training trajectory is constructed through complex pathways that warrant closer inspection. Concurrently, normalizing flows are experiencing a revival of interest as a classical family of likelihood-based methods. Recent advancements, such as TARFlow, have demonstrated that normalizing flows are fully capable of achieving competitive performance on complex image modeling tasks, positioning them as viable, mathematically rigorous alternatives to diffusion models and score-based approaches. Building upon these foundations, further developments such as iterative TARFlow seek to push the boundaries of normalizing flow generative models by incorporating iterative denoising strategies, aligning the trajectory-based benefits of modern diffusion paradigms with the exact likelihood estimation inherent to flow-based architectures. Post navigation Normalizing Trajectory Models Bridge the Gap Between Fast Generation and Exact Likelihood in Generative AI