← Back to the project page

Technical blog

How to Train a One-Step Language Model Without a Teacher

One- and few-step generative models for text usually get their speed by distilling a well-trained teacher model. For flow and diffusion language models, this requires a costly two-stage training pipeline. In our work, we show that one can train the endpoint map of a time-independent flow directly from data without a teacher, obtain clean sequences in one step, and refine the output with each additional NFE.

Based on Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning by
Sophia Tang and Shiyi Wang.

arXivalphaXivBibTeX

Authors

Sophia Tang

Affiliations

University of Pennsylvania
Harvard University

Published

Oct 9, 2026

Contents
  1. The challenge of few-step discrete generation
  2. Autonomous flows on the simplex
  3. Learning one-step maps without distillation
  4. Training and sampling with fixed-point refinement
  5. Experiments in language modeling and reasoning
  6. Conclusions
TL;DR Discrete Beckmann Transport Models (DBTM) introduce a new family of one-step language models that learn the endpoint map of a time-independent flow directly from data, with no teacher. Each extra function evaluation refines a complete sequence, achieving the Pareto frontier on LM1B and OWT and 84.6% and 16.8% accuracy on Sudoku and GSM8K in few NFE. [1]
Figure 1Transport map to the simplex in one-step. Gaussian samples $x_0\sim\mathcal N(0,\sigma^2 I)$ in the ambient space around the simplex are carried by a single transport map $T_\theta$ to the vertices $\{e_1,\dots,e_V\}^L$, which are clean tokens. The rest of this post explains where that map comes from and how to learn it.

Autoregressive language models generate one token at a time from left-to-right. Diffusion and flow models start from noise and denoise every position in parallel, which promises faster generation and bidirectional context. However, parallel decoding only matters if it significantly reduces the denoising steps needed to obtain a clean sequence, and that is exactly where discrete generative models struggle.

In this article, we will start with why few-step discrete generation is hard and why the usual fix, distilling a flow map from a teacher, is expensive. Then we look at what happens when a flow's velocity field no longer depends on time. Its trajectories end at the vertices of the simplex, and the map to those endpoints satisfies a simple conservation equation. That equation can be solved directly from data with an iterative optimization algorithm that is minimized at the time-independent one-step map at convergence. Finally, we turn the learned map into a sampler whose extra function evaluations act as refinement steps instead of ODE integration steps, and show that it excels on language modeling, Sudoku, and GSM8K against both continuous and discrete diffusion baselines in few NFE.

Notation $x$ is a state and $X_s(x)$ is the autonomous flow started from $x$, run for time $s$. $T(x)=X_{\tau(x)}(x)$ is the transport map associated with this flow, defined as the state the flow reaches at its hitting time $\tau(x)$ on the data. We reserve $t\in[0,1]$ for interpolation time and $s\ge 0$ for autonomous evolution time. For a sequence, $x^\ell$ is position $\ell$ and $x^\ell_j$ is vocabulary coordinate $j$. $\mu_0$ is the Gaussian prior, $\mu_1$ the data distribution, and $M_1=\operatorname{supp}(\mu_1)$ its support.
Part 01

The challenge of few-step discrete generation

Why parallel decoding needs many steps, and why compressing a flow usually needs a teacher.

1.1Parallel decoding breaks joint structure

Why does masked discrete diffusion need so many steps?

Consider a training set with exactly two sequences, New York and San Francisco, each with probability Β½. A masked diffusion model sees both positions masked and predicts each one from the same context. Its marginals are correct, since the first position is New or San with probability Β½ each, and the second is York or Francisco with probability Β½ each.

The problem appears when we sample both positions at once. Because each position is drawn independently, half the samples are New Francisco or San York, which never appeared in the data. With two steps, we unmask one position, condition on it, and only then predict the next, so every sample is valid. In general, discrete diffusion approximates each transition by a product over positions,

$$p_{t\mid s}(y_t\mid y_s)\;\approx\;\prod_{\ell=1}^{L}p^{\ell}_{t\mid s}\big(y^\ell_t\mid y_s\big),$$

which is nearly exact for tiny steps and increasingly wrong for large ones. Fewer steps means more tokens per step, which means more factorization error. No amount of model capacity fixes this, because a product of marginals cannot represent the correlation.

Figure 2Factorization error. Each masked position is predicted independently. Sampling both at once produces the mixtures New Francisco and San York half the time. Unmasking one position and conditioning on it before predicting the next recovers the joint structure.

1.2Generative flows and stochastic interpolants

A continuous-time path between a prior and the data distribution

Flow-based models avoid the factorization altogether. A stochastic interpolant [5] connects a sample from the prior to a sample from the data along an endpoint-conditioned path, and the model learns the average velocity of all paths passing through each point.

Interpolant $$\begin{gathered}x_t=I_t(x_0,x_1)=\alpha_t x_0+(1-\alpha_t)x_1\\ x_0\sim\mu_0,\quad x_1\sim\mu_1\end{gathered}$$
Marginal velocity $$b_t(x):=\mathbb E_{x_0,x_1}\big[\dot I_t\mid I_t=x\big]$$

Here $\alpha_0=1$ and $\alpha_1=0$, so $I_0=x_0$ and $I_1=x_1$. We use the schedule $\alpha_t=(1-t)^a$, where $a=1$ gives the linear interpolant. Each path $(x_0,x_1)$ has its own velocity $\dot I_t$. Many paths cross the same point at the same time, so the regression target $b_t(x)$ is their conditional mean. Integrating $\dot X_t=b_t(X_t)$ from $X_0\sim\mu_0$ reproduces the interpolant's marginals at every time, so $X_1\sim\mu_1$.

1.3Continuous flows on the simplex

How do we generate discrete sequences with a continuous flow?

Embed token $j$ as the one-hot vector $e_j\in\mathbb R^V$. The convex hull of these vectors is the probability simplex $\Delta^{V-1}$, and a sequence of $L$ tokens is a vertex configuration in $(\Delta^{V-1})^L$. The data distribution therefore lives on a very low-dimensional set, the vertices of the simplex.

Data manifold = vertices of the simplex $$M_1=\operatorname{supp}(\mu_1)=\big(\{e_1,\dots,e_V\}\big)^L\subset\big(\Delta^{V-1}\big)^L$$

We draw $x_0\sim\mathcal N(0,\sigma^2 I_V)$ independently at each position in the ambient space $\mathbb R^{V\times L}$ and flow toward the vertices. At the end, we read off each token with an argmax.

Figure 3A sequence is a point in a product of simplices. Each of the six positions starts from Gaussian noise and moves toward a vertex of its own simplex as $t$ goes from 0 to 1. An argmax at each position decodes the sentence.
Key idea Although the prior factorizes across positions, the flow that transports it does not. Every coordinate of $b_t$ can depend on the entire sequence state, so a deterministic joint transformation of independent noise can produce correlated tokens. Continuous flows do not need the factorization assumption.

1.4The cost of sampling with many function evaluations

Integrating the ODE requires many small steps

The catch is inference cost. To follow $b_t$ accurately, a flow-matching sampler needs many small Euler steps, and each step is one network evaluation (NFE). Reducing that count means taking fewer, larger steps, which requires knowing where a trajectory goes over a finite interval rather than only its instantaneous direction. A flow map provides exactly this by learning the solution operator, so it can take large jumps along the same trajectory.

Flow matching Β· many Euler steps $$\hat x_1=x_{t_0}+\sum_{i=0}^{N-1}b_{t_i}(x_{t_i})\,\Delta t$$
Flow maps Β· large jumps $$\hat x_1=X_{0.5,1}\big(X_{0,0.5}(x_0)\big)$$

1.5Flow maps for trajectory distillation

Jumping directly between any two times along the ODE

A flow map $X_{s,t}$ sends the state at time $s$ to the state at time $t$ on the same trajectory. It is convenient to write it in terms of a mean velocity over the interval [8].

Two-time flow map $$X_{s,t}(x_s)=x_t=x_s+(t-s)\,v_{s,t}(x_s),\qquad\forall\,s,t\in[0,1]$$

On the diagonal, $v_{s,s}=b_s$, so the instantaneous velocity can be trained with the usual flow-matching loss. Off the diagonal, there is no direct regression target. The model must learn how local predictions compose, which requires consistency objectives such as the semigroup identity $X_{t,u}\circ X_{s,t}=X_{s,u}$. Discrete targets add one more difficulty, because the posterior over token identity changes abruptly within a narrow window of $t$, which motivates the time reparameterization used by flow map language models [3].

1.6The distillation bottleneck

Few-step flow maps depend on a pretrained teacher and time conditioning

Flow maps make few-step generation possible, but the way they are usually trained introduces two costs.

  1. Few-step flow maps are distilled from a teacher flow. The student is capped at the teacher's quality and needs a two-stage pipeline. Self-distillation avoids the teacher by bootstrapping from the model's own predictions, but then it chases a moving target.
  2. Flow maps are conditioned on time. A two-time map $X_{s,t}$ has to learn every jump between every pair $(s,t)$, both diagonal and off-diagonal. That is a far larger space of targets than a sampler ever visits.
Figure 4Two pipelines. Flow-map distillation (top) trains a teacher flow $b_t$ to convergence, then distills $X_{s,t}$ from it. For FMLM on LM1B, that is a 1M-step teacher plus 100K distillation steps, and the student's quality is capped by the teacher. DBTM (bottom) trains a single transport map $T_\theta$ from data, with no teacher and no time input, in 200K steps on LM1B.

What we want is a one-step map that jumps to a clean sequence every time it is applied, and that we can learn from data alone.

The questionCan we learn a one-step map directly from data, with no teacher and no time conditioning?

Part 02

Autonomous flows on the simplex

Remove time from the velocity field, and what remains are fixed basins, absorbing vertices, and a map that is constant along every path.

2.1Taking the time average of the flow velocity

What if the velocity field did not depend on time?

Keep the same interpolant, but regress the velocity onto a field that sees only the state, with $t\sim\mathcal U[0,1]$ hidden from the model.

Time-dependent $$b_t(x):=\mathbb E_{x_0,x_1}\big[\dot I_t\mid I_t=x\big]$$
Autonomous Β· the same at every step $$b(x):=\mathbb E_{t,x_0,x_1}\big[\dot I_t\mid I_t=x\big]$$

The minimizer is still a conditional expectation, but now it averages over time as well as over endpoints. Write $\mu_t$ for the density of $I_t$. Because $t$ is uniform, a training state $x$ was produced at time $t$ with posterior probability $p(t\mid x)=\mu_t(x)/\int_0^1\mu_{t'}(x)\,dt'$. By the tower property,

Posterior-weighted time average $$b(x)=\int_0^1 b_t(x)\,p(t\mid x)\,dt=\frac{\int_0^1\mu_t(x)\,b_t(x)\,dt}{\int_0^1\mu_t(x)\,dt}.$$

Each time slice contributes where it is likely to have produced $x$. This is not a uniform average of the fields $b_t$. The resulting field defines an autonomous ODE, $\tfrac{d}{ds}X_s(x)=b(X_s(x))$ with $X_0(x)=x$. The autonomous flow runs on a separate clock (denoted $s\in [0,\infty)$) to the time-dependent flow and nothing forces it to reach the target at $s=1$, so it instead runs over a potentially infinite time horizon with hitting time $\tau(x)$.

What the time-dependent field does. For a single token with vertex probabilities $p_j$ and the linear schedule, the endpoint posterior is a softmax, $w_j(x,t)\propto p_j\exp\!\big(\beta(t)x_j\big)$ with $\beta(t)=t/\big(\sigma^2(1-t)^2\big)$. The field is $b_t(x)=\big(D_t(x)-x\big)/(1-t)$, where $D_t=\sum_j w_j e_j$ is the posterior-mean denoiser. If we freeze $t$, the equilibria solve $x=\operatorname{softmax}(\log p+\beta(t)x)$. At $t=0$ there is a single attractor at the mean $p$. As $\beta$ grows, new attractors appear at staggered times and move from the interior toward the vertices. A particle chasing moving attractors follows a curved path.

What the autonomous field does. Its stable equilibria are exactly the vertices, and its basins of attraction never move.

Figure 5Moving versus fixed destinations. On the left, the time-dependent field $b_t$ evolves as $t$ advances, and its attractors appear at staggered times inside the simplex before converging on the vertices. On the right, the autonomous field $b$ is stationary, with stable equilibria only at the vertices.
Freeze the clock and find the attractorsInteractive 1 Β· computed live
interpolation time $t$0.30

time-dependent $b_t(x)$

frozen at t = 0.30

autonomous $b(x)$

the same for every $t$

stable equilibrium of the frozen fieldsimplex $\Delta^2$

Both fields are computed exactly in your browser for a single token with $V=3$, $p=(0.45,0.33,0.22)$, $\sigma=0.55$, and a linear interpolant. They are shown on the invariant plane $\sum_jx_j=1$. The autonomous field is the posterior-weighted time integral above, evaluated by quadrature. Sweep $t$ and watch the frozen field's attractors split and migrate while the autonomous panel stays put.

2.2Beckmann's transportation problem

Does the autonomous flow provably generate the target distribution?

An autonomous flow does not reproduce the interpolant's marginals at intermediate times, but it inherits a conservation law from them. Integrating the continuity equation $\partial_t\mu_t+\nabla\cdot(b_t\mu_t)=0$ over $t\in[0,1]$ telescopes the time derivative into the two endpoint densities.

Continuity equation β†’ divergence condition $$\int_0^1\nabla\cdot(b_t\mu_t)\,dt=-\int_0^1\partial_t\mu_t\,dt\quad\Longrightarrow\quad\nabla\cdot j=\mu_0-\mu_1$$
Current $$j=\nu\,b=\int_0^1 b_t\,\mu_t\,dt$$
Occupation measure $$\nu(x):=\int_0^1\mu_t(x)\,dt$$

The prior supplies probability mass into the time-integrated current $j$ (source), and the data manifold becomes the sink that removes it. Because $\nu\gt0$ away from the data, $j$ and $b$ point in the same direction. They share oriented flow lines and differ only in speed, so the autonomous flow moves mass along the lines of a current that starts at $\mu_0$ and ends at $\mu_1$. When $\mu_1$ is singular, as it is for discrete data, the equality holds in the distributional sense. This source–sink constraint is the one in Beckmann's transportation problem [6], which gives the method its name, and its connection to Beckmann's problem and to optimal transport is developed in the Beckmann Transport Models paper [2].

The divergence condition alone does not rule out mass escaping to infinity or stalling in the interior. Under conditions on the prior, the target support, and the interpolant, the BTM analysis shows that the autonomous flow reaches the target support in finite time and that its endpoint map pushes the prior onto the target, $T_\#\mu_0=\mu_1$ [2, Prop. 1]. The next section states the corresponding result on the simplex.

Figure 6The autonomous flow. Prior samples follow the stationary field $b$ (dashed flow lines) and collect at the three vertices. The clock $s$ runs past 1, since each trajectory takes as long as it needs.

2.3Every path ends on a vertex

Convergence of the autonomous flow on the vertices of the simplex in finite time

The divergence identity alone does not rule out probability escaping to infinity or getting trapped in the interior. Our convergence theorem closes those gaps for the simplex [1, Thm. 3.1], building on the continuous-state result for Beckmann transport models [2]. For almost every starting point $x_0$, the following three properties hold.

(i) Absorption $$X_s(x_0)=X_{s^\star}(x_0)\quad\text{for } s\gt s^\star$$
(ii) Convergence $$\lim_{s\to\infty}X_s(x_0)=x^\star(x_0)\in M_1$$
(iii) Finite hitting time $$\lim_{r\downarrow0}b(e_j+r\omega)=-\kappa_j(\omega)\,\omega$$

Part (iii) proves a unique feature of the autonomous flow of the linear interpolant. The field vanishes at the vertex, but its limit along every incoming direction $\omega$ has speed $\kappa_j(\omega)\ge\kappa_-\gt0$. Writing $h(s)=\|X_s-e_j\|$ for the distance to the vertex, once a trajectory enters the ball $B(e_j,r_1)$, where $r_1$ is small enough that $\langle b(e_j+r\omega),\omega\rangle\le-\kappa_-/2$ for all $0\lt r\le r_1$, we get

$$\dot h(s)=\Big\langle b(X_s),\,\tfrac{X_s-e_j}{h(s)}\Big\rangle\le-\frac{\kappa_-}{2},$$

so from any time $s$ at which the trajectory is inside the ball, it reaches $e_j$ by $\tau\le s+2h(s)/\kappa_-\le s+2r_1/\kappa_-$. The particle arrives with nonzero speed and then stops at finite time.

What happens exactly at a vertex. The limit above is not the value at the vertex. If $\alpha_t\sim(1-t)^a$ as $t\to1$ with $a\ge1/V$ and the prior decays exponentially, the autonomous field vanishes at every vertex, $b(e_j)=0$, so absorbed trajectories stay put [1, Thm. C.1]. For the linear interpolant the field is therefore discontinuous at the vertices, vanishing at $e_j$ but approaching it with speed at least $\kappa_-$. This discontinuity is what makes arrival take finite time. If the schedule instead slows to a stop, $\dot\alpha_1=0$, then $b(x)\to0$ as $x\to e_j$ and the field is continuous at the vertices [1, Prop. C.5]. The speed bound above is specific to the linear case.

Why no mass escapes or stalls. We prove two lemmas that justify why no mass escapes or stalls away from the vertices [1, Lemmas C.4–C.5]. Far from the simplex the field behaves like $b(x)=-x\,(1+o(1))$, so $\langle b(x),x\rangle\lt0$ outside some ball $B(0,R_0)$, so every trajectory enters the ball in finite time and never leaves it. Inside the ball, the divergence condition away from the data, $\nabla\cdot(\nu b)=\mu_0$, gives $\tfrac{d}{ds}\big[\nu(X_s)\,J_s\big]=\mu_0(X_s)\,J_s$, where $J_s$ is the Jacobian determinant of the flow. A set of starting points that never approaches the data would carry an occupation mass that grows exponentially, by GrΓΆnwall's inequality, while staying bounded by $\nu$ of a bounded region. So that set has Lebesgue measure zero (it occupies no volume). Combined with finite arrival near the vertices, this gives the three properties of Theorem 3.1.

Figure 7The three parts of the theorem. (i) A trajectory reaches its vertex at $s^\star$ and stays there. (ii) Every prior sample converges to a vertex in the long-time limit (300/300 in this toy). (iii) Near $e_1$ the field points inward from every direction, with limiting speed at least $\kappa_-$, and vanishes only at the vertex.
Figure 8Each path stops at its own hitting time $\tau(x_0)\lt\infty$. Three trajectories reach their vertices at $\tau\approx1.12$, $1.20$, and $1.22$ and remain there.

2.4Autonomous flow has straighter trajectories

Why prefer the autonomous flow over the time-dependent one?

A time-dependent trajectory curves because its field changes underneath it. It heads toward the mean, then bends as attractors appear and move. The autonomous field's basins are fixed from the start, so each particle heads almost directly to its destination. In the paper's toy comparison on the 2-simplex, the mean ratio of path length to chord length is about 1.09 for the autonomous flow versus 1.58 for the time-dependent flow. Straighter paths are easier to compress into a single jump.

Figure 9Path length / chord length. Autonomous trajectories (solid) versus time-dependent trajectories (dashed) from the same starting points. In this animation's sample of 153 prior points, the mean ratios are 1.042 and 1.294.
Drop a point and compare the time-dependent and autonomous flowsInteractive 2 Β· computed live
autonomous $\dot X=b(X)$time-dependent $\dot X=b_t(X)$basins of $e_1, e_2, e_3$
autonomous path / chordβ€”
time-dependent path / chordβ€”
trajectories0

Click or tap anywhere on the plane to start a trajectory from that point.

The same toy as Interactive 1. Each start point is integrated under both fields, the autonomous flow until it is absorbed at a vertex, and the time-dependent flow from $t=0$ to $1$. The ratios are averaged over the trajectories on screen. The basin coloring shows where the autonomous flow sends each point, which is the transport map $T$ of Section 2.6.

2.5Eulerian and conservation equations

The defining equations of the autonomous flow and its transport map

Autonomous flows have a semigroup property, meaning that following the flow for time $\varepsilon$ and then for time $s$ is the same as following it for $s+\varepsilon$. Compare two ways of getting there. We can run the flow from $x$ a little longer, $X_{s+\varepsilon}(x)\approx X_s(x)+\varepsilon\,\partial_sX_s(x)$. Or we can first nudge the starting point along the field and then run for time $s$, $X_s(x+\varepsilon b)\approx X_s(x)+\varepsilon\,b\cdot\nabla X_s(x)$. Both land on the same point, which gives an equation for the flow as a function of its starting point.

Eulerian equation $$\begin{cases}b(x)\cdot\nabla X_s(x)=\partial_sX_s(x), & x\notin M_1,\\ X_s(x)=x, & x\in M_1.\end{cases}$$
Figure 10Deriving the Eulerian equation. Running the flow a little longer from $x$ (step (2)) and nudging the start along $b$ before running (step (3)) reach the same point. Equating the two expansions gives $b\cdot\nabla X_s=\partial_sX_s$. Use the chapter buttons to step through.

Now let $s$ approach the hitting time. Once the trajectory is absorbed, it stops moving, so the right-hand side vanishes, with $\partial_sX_s(x)\to0$ as $s\to\tau(x)$. The limit of $X_s$ is the endpoint map $T$, and it satisfies a stationary equation.

Conservation equation $$\begin{cases}b(x)\cdot\nabla T(x)=0, & x\notin M_1,\\ T(x)=x, & x\in M_1.\end{cases}$$

The equation says that $T$ does not change as you move along the flow. Every point on the same path ends at the same vertex, so $T$ is constant along each trajectory. Moving across the flow, at a basin boundary, $T$ can jump. The exact endpoint map is piecewise constant, with discontinuities on basin boundaries.

Figure 11Reading the conservation equation. (1) $T(x)$ is where the flow from $x$ ends. (2) Sliding $x$ along the flow never moves the endpoint, so $b\cdot\nabla T=0$. (3) Sliding $x$ across the flow can make the endpoint jump, since $\nabla T\ne0$ but $\nabla T\perp b$. (4) At a vertex nothing moves, so $T(x)=x$ on $M_1$.

2.6The autonomous transport map

Can we skip the path and jump straight to its end?

The one-step transport map is the long-time limit of the autonomous flow.

One-step transport map $$T(x_0):=\lim_{s\to\tau(x_0)}X_s(x_0)\in M_1,\qquad T_\#\mu_0=\mu_1$$

The long-time limit is on the left and the pushforward is on the right, meaning $\mu_0\big(T^{-1}(A)\big)=\mu_1(A)$ for every set $A$.

It has three properties that the rest of the method relies on.

  1. Constant along every trajectory. $T\big(X_s(x_0)\big)=T(x_0)$.
  2. Idempotent. $T\circ T=T$, so $T^{\circ k}=T^{\circ k'}$ for all $k,k'\in\mathbb N$. Once you reach the data, applying $T$ again changes nothing.
  3. Its level sets are the basins of the vertices.

The boundary condition is a key component of the definition. Any constant function $T(x)=c$ solves $b\cdot\nabla T=0$, but no constant function fixes all of the vertices. If $\widetilde T$ is constant along the same trajectories and equals the identity on $M_1$, then $\widetilde T(x)=\lim_{s\uparrow\tau(x)}\widetilde T(X_s(x))=T(x)$. The interior equation and the boundary condition together define the endpoint map uniquely along the trajectories that reach the data [1, Prop. 3.2].

Figure 12The transport map partitions space into basins. Every prior sample is sent to the vertex whose basin contains it (lavender $e_1$, sky $e_2$, mint $e_3$). Applying the map again leaves it in place, so $T(T(x_0))=T(x_0)$. Try the basin toggle in Interactive 2 to compute these regions yourself.
Part 03

Learning one-step maps without distillation

Turn the conservation equation into a regression problem that needs only data samples.

3.1Iteratively solving the conservation equation

How do we train $T$ without a teacher?

We use the Eulerian and conservation equations to define a training algorithm whose objective is minimized at the transport map. Initializing $T=\mathrm{id}$ as the identity (which already satisfies the boundary condition) we can iteratively minimize the residual of the conservation equation which pushes the map along the autonomous field until the residual vanishes when $T$ is the transport map.

Euler iteration of step size $\eta$ $$X^{(n+1)}(x)=X^{(n)}(x)+\eta\,b(x)\cdot\nabla X^{(n)}(x),\qquad X^{(0)}=\mathrm{id},\qquad X^{(n)}\to T$$

Each training iteration advances the flow a little further, and the fixed point satisfies $b\cdot\nabla T=0$. One problem remains, which is that we don't have access to the marginal autonomous velocity $b$ (a quantity often estimated with a teacher model). However, since $b(x)$ is the conditional mean of the sampled velocity $\dot I_t$, and the directional derivative is linear in the direction, we can replace $b$ by $\dot I_t$ inside a regression without changing the fixed point.

$$\mathbb E_{t,x_0,x_1}\big[\dot I_t\cdot\nabla T_\theta(I_t)\mid I_t=x\big]=b(x)\cdot\nabla T_\theta(x).$$

Regressing $T_\theta(I_t)$ onto the detached target $T_\theta(I_t)+\dot I_t\cdot\nabla T_\theta(I_t)$ performs one Euler step in expectation. Adding the boundary condition at clean data gives the optimal map as the solution of the following objective:

Self-regression target + boundary condition $$T=\arg\min_{T_\theta}\;\mathbb E_{t,x_0,x_1}\Big[\big\|T_\theta(I_t)-\operatorname{sg}\big(T_\theta(I_t)+\dot I_t\cdot\nabla T_\theta(I_t)\big)\big\|^2+\big\|T_\theta(x_1)-x_1\big\|^2\Big]$$

The directional derivative is a single Jacobian–vector product, so training does not need the marginal velocity from a pretrained teacher.

Figure 13From identity to transport map. At $n=0$, the map is the identity and every point stays put. Each iteration pushes $X^{(n)}(x)$ further along the flow. For large $n$, the iterates converge to $T$, which sends every point to its vertex.

3.2A partially trained map still converges

What happens if training the map is not converged?

Suppose the trained map does not converge. The iterative training scheme means the map we have learned approximates the autonomous flow truncated at some finite time $T_\theta\approx X_t$, which can yield better and better approximations of the true map $T$ when we apply it repeatedly by the semigroup property.

Truncated flow $$T_\theta(x_0)\approx X_t(x_0)$$
Semigroup $$X_t\circ X_t=X_{2t}$$
Iterating the map $$T_\theta^{\circ k}(x_0)\approx X_{kt}(x_0)\;\longrightarrow\;M_1\quad\text{as } kt\to\tau(x_0)$$

Reapplying the map $k$ times acts like jumping along the autonomous flow, and the iterates stop changing once $kt$ reaches the hitting time. Generation therefore reduces to iterating one map until it reaches a fixed point.

Figure 14Iterating a partially trained map. Each application $T_\theta^{\circ k}(x_0)\approx X_{kt}(x_0)$ advances a point along its trajectory. After a few applications, it sits at a vertex and stays there.

3.3Training objectives

How to train the transport map effectively from data

For sequences, we parameterize $T_\theta(\cdot)=\operatorname{softmax}(f_\theta(\cdot))$ which outputs a categorical distribution at each position. The sums below run over positions outside a clean context set $\mathcal C$, which is introduced in Section 4.1 and is empty when no clean context exists before the first NFE. We use the following objectives, each of which encodes a constraint satisfied by the transport map.

Euler self-regression step $$\mathcal L_{\text{transport}}(\theta):=\mathbb E_{t,x_0,x_1}\Big[\sum_{\ell\in[L]\setminus\mathcal C}\big\|T_\theta(I_t)^\ell-\operatorname{sg}\big(T_\theta(I_t)+\dot I_t\cdot\nabla T_\theta(I_t)\big)^\ell\big\|^2\Big]$$
Boundary condition at clean data, $T(x)=x$ for $x\in M_1$ $$\mathcal L_{\text{bnd}}(\theta):=\mathbb E_{x_1}\Big[\sum_{\ell\in[L]\setminus\mathcal C}\operatorname{CE}\big(T_\theta(x_1)^\ell,\,x_1^\ell\big)\Big]$$
Semigroup condition $$\mathcal L_{\text{semi}}(\theta):=\mathbb E_{t,x_0,x_1}\Big[\sum_{k\in\mathcal K}\sum_{\ell\in[L]\setminus\mathcal C}\operatorname{CE}\big(T_\theta(I_t)^\ell,\,\operatorname{sg}\big(T_\theta^{\circ k}(I_t)\big)^\ell\big)\Big]$$
Full objective $$\mathcal L(\theta)=\mathcal L_{\text{transport}}+\lambda_{\text{semi}}\mathcal L_{\text{semi}}+\lambda_{\text{bnd}}\mathcal L_{\text{bnd}}+\lambda_{\text{anchor}}\mathcal L_{\text{anchor}}$$

The Euler-corrected target does not necessarily lie on the simplex, so the transport term uses squared error, while clean categorical targets use cross-entropy. The semigroup term is only effective late in training, because the identity map is already idempotent and satisfies the constraint, so consistency alone does not yield the transport we want. That leaves one more question. From our previous discussion, we notice that the boundary term only supervises perfectly clean inputs whereas all intermediate points along a flow should land on the vertices. While the transport objective enforces consistency over the trajectory, it is a rather indirect objective which preserves the posterior at any point. However, there is a point along each interpolant where the posterior reduces to a Dirac delta at a single vertex, which motivates the question: When and how can we directly enforce that points that have already committed to a vertex land on it?

3.4Commitment time of the flow trajectory

When does the time-dependent flow commit to a vertex on the simplex?

A naive way to anchor all ambient points to land on a vertex is to minimize the discrepancy between $T(I_t)$ and $x_1$ for all $t\in[0,1]$. However, this constraint only holds when the target is fixed and the posterior at $I_t$ is $\delta(x_1)$, and it does not hold for general times $t$ when each position must resolve its token identity relative to the other tokens in the sequence to form a jointly coherent sequence. Only after a specific time along the interpolant are all the tokens committed to a specific target, and only then can we enforce an additional loss that anchors to the committed target. We call this the commitment time $t^*$, and we derive an expression for it next.

A signal-plus-noise problem. Take one token with a uniform prior over $V$ vertices, fix its true endpoint $e_k$, and write $x=I_t=(1-\alpha_t)e_k+\alpha_tz$ with $z\sim\mathcal N(0,\sigma^2I)$ and $\alpha_t=(1-t)^a$. The posterior logits of the sampled state are

Planted logits $$U_j=\rho^2\,\mathbf 1[j=k]+h_j,\qquad h_j\overset{\text{iid}}{\sim}\mathcal N(0,\rho^2),\qquad \rho(t)=\frac{1-\alpha_t}{\alpha_t\,\sigma}.$$

Everything about the schedule and the noise scale has collapsed into one signal-to-noise ratio $\rho$. The true token receives a planted signal $\rho^2$, and every token receives Gaussian noise with standard deviation $\rho$. The posterior mass on the target is

$$p(k\mid x)=\frac{e^{\rho^2+h_k}}{e^{\rho^2+h_k}+\sum_{j\ne k}e^{h_j}},$$

and the competing sum has the form of a partition function in Derrida's Random Energy Model [7], with $V-1$ independent random energies. The typical largest competitor is $\max_{j\ne k}h_j\approx\rho\sqrt{2\log V}$. The target wins once its signal beats that maximum.

$$\underbrace{\rho^2}_{\text{target signal}}\approx\underbrace{\rho\sqrt{2\log V}}_{\text{largest competitor}}\quad\Longrightarrow\quad\rho^\star\approx\sqrt{2\log V}.$$

Because $\rho(t)$ increases monotonically in $t$, we can invert this threshold to get a time.

Phase transition / commitment time [1, Thm. 4.1] $$t^*=1-\Big(1+\sigma\sqrt{2\log V}\Big)^{-1/a}$$

Before $t^*$, the token is undecided and must be resolved jointly with the other positions. After $t^*$, the target vertex dominates, and further integration does not change the output. Commitment time depends on the vocabulary size $V$, the prior scale $\sigma$, and the schedule exponent $a$. Larger noise delays it, a faster-decaying schedule moves it earlier, and $V$ enters only through a logarithm.

Figure 15Crossing $t^*$. One interpolant on the 2-simplex ($V=3$, $\sigma=0.55$, $a=1$, so $t^*\approx0.45$). The posterior $p(x_1=e_k\mid I_t)$ moves gradually at first, then snaps to the target once the path enters the anchor region past $t^*$.
Simulate when the target vertex beats $V-1$ competing verticesInteractive 3 Β· simulated live
interpolation time $t$0.60
noise scale $\sigma$1.00
schedule exponent $a$1.0
vocabulary
signal-to-noise $\rho(t)$β€”
competitors above targetβ€”
posterior $p(k\mid x)$β€”
commitment time $t^*$β€”

One draw of $z$ is sampled per vocabulary and reused for every $t$, so moving the slider follows a single interpolant. The histogram shows competitor logits in units of $\rho$. The dashed line marks the predicted maximum $\sqrt{2\log V}$, and the indigo spike marks the target at $\rho+z_k/\sigma$. The lower curve is the exact posterior of the target for this draw, with $t^*$ from the formula marked in mint.

3.5Commitment time across vocabularies

How does $t^*$ scale with vocabulary size and schedule?

The table below evaluates the commitment time for the vocabularies used in our experiments, with $\sigma=1$ and $a=1$.

LM1BOpenWebTextGSM8KSudoku
Vocabulary $|V|$30,52250,25749,15212
Noise scale $\sigma$ Β· exponent $a$1 Β· 11 Β· 11 Β· 11 Β· 1
$\sqrt{2\log|V|}$4.5454.6534.6482.229
Commitment time $t^*$0.8200.8230.8230.690

Because $V$ enters only through $\sqrt{2\log V}$, the language vocabularies all commit at nearly the same time, while a 12-token Sudoku vocabulary commits much earlier. The commitment time gives us a principled answer to the question at the end of Section 3.3. Over $[t^*,1]$ the endpoint has already been decided, so we can anchor the map to it directly.

Anchor loss Β· enforces $T_\theta(I_t)=x_1$ for $t\gt t^*$ $$\mathcal L_{\text{anchor}}(\theta):=\mathbb E_{t,x_0,x_1}\Big[\sum_\ell\mathbf 1[t\gt t_{\text{anchor}}]\,\operatorname{CE}\big(T_\theta(I_t)^\ell,\,x_1^\ell\big)\Big]$$

Before commitment, the transport objective shapes how the map changes along the flow and leaves different starting states free to choose different endpoints. After commitment, the anchor supplies a direct target. This also replaces the delicate time reparameterization that flow-map language models need. We still sample $t$ during training, but the network itself never sees it.

The commitment times for each datasetInteractive 4
noise scale $\sigma$1.00
custom vocabulary $V$1,000
$a=1$$a=2$$a=3$datasetsyour $V$

Computed from $t^*=1-(1+\sigma\sqrt{2\log V})^{-1/a}$. Increasing $\sigma$ pushes every curve later, and a larger exponent $a$ moves commitment earlier. The custom-$V$ marker reads off all three schedules at once.

We found that changing the anchor time has a real impact empirically. Anchoring too early supervises toward a fixed target while the endpoint is still undecided, and the model collapses onto repetitive tokens while anchoring too late leaves the post-transition interval unsupervised and slows convergence [1, Table 12].

Figure 16Anchor-time ablation on LM1B (DBTM + ril, 16 NFE, $\kappa=0.9$). Anchoring at $0.60$ yields a deceptively low Gen-PPL of 26.0, but entropy collapses to 3.84. Anchoring at $0.85$ preserves entropy but raises Gen-PPL to 103.5. The default $t_{\text{anchor}}=0.75$ sits between the two failure modes.
Part 04

Training and sampling with fixed-point refinement

Because every application proposes a clean sequence, extra NFEs can refine a proposal rather than integrate an ODE.

In the exact theory, $T\circ T=T$, so one application reaches the data and further applications do nothing. A learned $T_\theta$ is only an approximation, and its first proposal can contain inconsistencies. While flow maps use additional computation to take smaller jumps along the ODE, DBTM generates the fixed endpoint by construction, so additional computation can be interpreted as fixed-point iteration that refines the final output.

4.1Clean-context interpolant

Training the map on partially clean sequences so that NFEs scale as refinement steps

Let $\mathcal C$ be a set of positions held at their clean values. We extend the interpolant so that context positions are one-hot tokens and every other position is noised to a single shared time $t$.

Clean-context interpolant $$I_t^\ell(\mathcal C):=\begin{cases}x_1^\ell, & \ell\in\mathcal C\quad\text{(one-hot tokens at context positions)},\\ \alpha_tx_0^\ell+(1-\alpha_t)x_1^\ell, & \ell\in[L]\setminus\mathcal C\quad\text{(noised to a shared time } t).\end{cases}$$

$T$ stays time-independent, but it is trained on this larger family, with a fixed fraction of samples that have no context at all. At inference, committed tokens become clean context, and the same map completes the rest.

Figure 17Training on partial context. Each draw holds a random context set $\mathcal C$ at $x_1$ (highlighted) and noises all other positions to the same $t$. The map never sees $t$.

4.2Commit, renoise, refine

Each NFE refines a partially clean sequence rather than integrating an ODE

The inference-time refinement loops through three steps.

  1. Propose and score. Apply the map, decode $\hat x^\ell=\arg\max T_\theta(\hat x)^\ell$, and compute a quality score $q^\ell$ for each token. Create the commit set with $q^\ell\ge\kappa$.
  2. Commit and renoise. Hold committed positions at their decoded tokens, $\hat x^\ell\leftarrow\arg\max(T_\theta(\hat x)^\ell)$ for $\ell\in\mathcal C$, and draw fresh noise $\hat x^\ell\sim\mathcal N(0,\sigma^2 I_V)$ everywhere else.
  3. Refine. Apply the map again, $\hat x\leftarrow T_\theta(\hat x)$, until every position is committed.

Each round costs one NFE, and the quality score comes from the same forward pass. Commitment also gives a natural stopping rule: sampling ends when every token passes the threshold, so the number of NFEs adapts to the sample instead of being fixed in advance.

Figure 18Four rounds of commit, renoise, refine (illustrative schedule). Positions with $q^\ell\ge\kappa=0.9$ are committed (shaded). The rest are renoised and proposed again with more context, until all twelve positions are committed.

4.3Commitment rules

Confidence- and quality-based commitment without extra NFEs

Each refinement round computes a per-token score that determines which positions are reliable enough to commit, and we consider two ways to compute it from the same forward pass with no extra NFEs.

Confidence-based score from the softmax probability of the argmax token $$q^\ell\big(T_\theta(x)\big):=\big\langle T_\theta(x)^\ell,\,\hat x_1^\ell\big\rangle,\qquad\hat x_1=\arg\max\big(T_\theta(x)\big)$$
Quality-based score from the learned probability of being correct given the context $$q^\ell_\phi\big(T_\theta(x)\big)=\Pr\big[\hat x_1^\ell=x_1^\ell\mid T_\theta(x)\big]$$

The quality head shares the network trunk and is trained by binary cross-entropy against whether each proposed token is correct.

$$\mathcal L_\phi=\mathbb E_{x_0,x_1}\Big[\sum_\ell\operatorname{BCE}\Big(\mathbf 1\big[\hat x_1^\ell=x_1^\ell\big],\,q^\ell_\phi\big(T_\theta(x)\big)\Big)\Big].$$

A threshold alone could stall if nothing is confident. Under a budget of $k$ rounds, with $\mathcal R_r$ the unresolved positions entering round $r$, we also ensure a minimum set of highest scoring tokens are committed.

Threshold βˆͺ commitment floor $$n_r=\Big\lceil\tfrac{|\mathcal R_r|}{k-r+1}\Big\rceil,\qquad\Delta\mathcal C_r=\{\ell\in\mathcal R_r:q^\ell_r\ge\kappa\}\;\cup\;\operatorname{top}_{n_r}(\mathcal R_r;q_r).$$

The threshold keeps confident positions, and the commitment floor guarantees progress. In the last round, $n_k=|\mathcal R_k|$, so every remaining position is committed within $k$ NFEs.

Run the commit rule yourselfInteractive 5
threshold $\kappa$0.90
budget $k$

committed (clean context)score $q^\ell_r$ of an uncommitted positioncommitted this round$\kappa$

Illustration of the renoise-refine loop. The scores are simulated rather than model outputs, and each round the uncommitted positions draw new scores whose typical value rises with the fraction of context already committed. The commit rule itself is exactly the threshold βˆͺ commitment floor rule above. Lower $\kappa$ or a smaller budget $k$ to see the commitment floor take effect.

4.4Closing the train–inference gap

Random training contexts β‰  the structured contexts the sampler builds

During standard training, context positions are chosen at random. During sampling, the context is chosen by the model, depends on its confidence, and can contain errors. This results in exposure bias, where the contexts the model builds from its own predictions at inference are out of distribution for the model. Removing time conditioning does not remove this bias. Refinement-in-loop (ril) training therefore constructs the partial-context interpolant from real, stop-gradient sampling rollouts of the current model and trains with the intermediate states of the rollout.

Rollout with confidence-based sampling, committing ground-truth tokens $$\mathcal L_{\text{ril-c}}(\theta):=\mathbb E_{x_0,x_1}\Big[\sum_{k\in\mathcal K}\sum_{r=1}^k\sum_{\ell\in[L]\setminus\mathcal C_r}\operatorname{CE}\big(T_\theta(\hat x_{r-1,k})^\ell,\,x_1^\ell\big)\Big]$$
Matching deeper rollouts after a warm-up phase $$\mathcal L_{\text{semi-r}}(\theta):=\mathbb E_{x_0,x_1}\Big[\sum_{k\in\mathcal K}\sum_{r=1}^k\sum_\ell\operatorname{CE}\big(T_\theta(\hat x_{r-1,k})^\ell,\,\operatorname{sg}\big(T_\theta(\hat x_{k-1,k})\big)^\ell\big)\Big]$$

The right target depends on the coupling. In Sudoku, the clues determine the solution, so the rollout target can stay the ground-truth grid. In unconditional language generation, an independently paired noise sample and training sequence do not define a unique correct trajectory. If a rollout is already forming one coherent sentence, forcing it toward an unrelated paired sentence works against the model's own transport assignment. After a warm-up phase, we therefore distill earlier rounds toward the detached output of a deeper rollout that is assumed to be more coherent, which supplies a target that is consistent with the rollout's own noise-to-output coupling. This provides a way to further enhance sequence quality in the few-step regime but requires the map to be trained well on data first.

4.5Training with refinement-in-loop

Refinement-in-loop training improves convergence on both language modeling and reasoning

Turning on refinement-in-loop (ril) training partway through shows its effect directly on both language modeling and reasoning.

Validation generative perplexity and entropy on LM1B at NFE 1, 2 and 4, before and after enabling refinement-in-loop training at 100K steps.
Validation Sudoku solve accuracy for DBTM and DBTM with refinement-in-loop training on easy, medium and hard puzzles.
Figure 19Turning on ril. At the top, on LM1B, Gen-PPL drops after the ril gate at 100K steps (dashed line) while entropy stays roughly constant. The committed variant (violet) improves more than the non-committed one (blue). At the bottom, on Sudoku, ril dramatically speeds up convergence. Best solve accuracy improves by +0.4, +3.5, and +19.9 points on easy, medium, and hard.

Refinement-in-loop improves even the first proposal (1 NFE). Relative to base DBTM, one-step generative perplexity drops by about 58% on LM1B and 65% on OpenWebText.

Figure 201-NFE Gen-PPL ↓. On LM1B it falls from 201.2 to 85.5 (βˆ’58%), and on OWT from 188.2 to 65.2 (βˆ’65%). Entropies are 4.03 β†’ 3.92 and 4.97 β†’ 4.95 (Table 1).
Part 05

Experiments in language modeling and reasoning

Toy dynamics, unconditional generation on LM1B and OpenWebText, and conditional reasoning on Sudoku and GSM8K.

5.1Dynamics of the time-dependent flow on the simplex

The attractors of the time-dependent flow appear at late times

We first conduct toy experiments to characterize the dynamics of the time-dependent flow on the simplex. On a 3-simplex with a near-tie between the two most likely tokens, $p=(0.40,0.38,0.20,0.02)$, only one attractor exists from $t=0$ to $t=0.5$, at the dominant vertex. The others appear one at a time after $t=0.5$, at $0.535$, $0.607$, and $0.687$. Even a vertex with 95% of the leader's probability does not appear until $t=0.535$. This is why time-dependent trajectories first head toward the dominant vertex and then bend, and why we prefer the autonomous field, where all attractors exist at the vertices with defined probability basins from the start.

When do the other attractors form? Early on, $x_t$ is mostly noise, so the flow falls back on the prior. A single attractor starts at the mean $p$ and drifts toward the most likely vertex $e_l$. As $t$ grows, $\beta(t)=t/\big(\sigma^2(1-t)^2\big)$ acts as a rising inverse temperature. Writing $x_j$ for the $j$-th coordinate of a fixed point $x$ of the frozen field (its weight on vertex $e_j$), dividing the fixed-point equations for coordinates $l$ and $j$ cancels the softmax normalization,

$$\frac{x_l}{x_j}=\frac{p_l}{p_j}e^{\beta(t)\,(x_l-x_j)},$$

where the prior favors $e_l$ and the exponential favors whichever coordinate is already larger, and that pull strengthens as $\beta$ grows. The attractor near $e_j$ appears when this equation first gains a stable solution, which happens approximately when

$$\beta-\log\beta\;\simeq\;1+\log\frac{p_l}{p_j}+\log(1+m),\qquad\text{i.e.}\qquad\frac{t_j}{(1-t_j)^2}\simeq-\sigma^2\,W_{-1}\!\left(-\frac{p_j}{p_l\,(1+m)\,e}\right),$$

where $W_{-1}$ is the lower branch of the Lambert $W$ function and $m$ counts the probabilities comparable to but less than $p_l$ [1, Props. C.2–C.4, Cor. C.1]. Attractors therefore form in order of probability: if $p_j\gt p_k$, then $t_j\lt t_k$. For a standard Gaussian prior ($\sigma=1$), no attractor other than the dominant one forms before $t=0.5$. For the example below ($m=1$), the formula gives $0.553$, $0.598$, and $0.679$, against the exact times $0.535$, $0.607$, and $0.687$.

Figure 21Even a near-tie appears late. The frozen time-dependent field on $\Delta^3$ with $p=(0.40,0.38,0.20,0.02)$ and $\sigma^2=1$, computed as $t$ sweeps from 0.3 to 0.8. Dots are uniform samples of the simplex colored by the attractor they converge to, and large markers are the attractors. Only the dominant attractor exists before $t=0.5$, and the others appear one at a time, at $t_2=0.535$, $t_3=0.607$, and $t_4=0.687$, even though $p_2/p_1=0.95$.

5.2Unconditional language modeling at few NFEs

DBTM reaches the coherence–diversity Pareto frontier with a few refinements

We train on LM1B and OpenWebText (OWT) and report generative perplexity (Gen-PPL) under GPT-2 Large [9] together with unigram entropy. Baselines include distilled discrete diffusion and continuous flow maps evaluated at matched NFE budgets. Both DBTM variants train for 200K steps, while FMLM is distilled for 100K steps from a 1M-step teacher.

Explore Table 1Interactive 6
DBTM (ours)continuous flow mapsdiscrete diffusion / othercollapsed (entropy far below the data)

Bars show Gen-PPL on a log scale (shorter is better), and the gray number is entropy. A low perplexity only counts if entropy stays near the data's.

All values from Table 1 of the paper [1]. Note that for DFM (ESD) on OWT at 1 NFE, the Gen-PPL of 5.33 comes with an entropy of 0.26, which means repeated tokens. That is why we compare perplexities only at entropies near the data and report full frontiers below.

A single perplexity number can be misleading, because a model that repeats tokens scores deceptively well. We therefore compare full Pareto frontiers over every inference-time knob (commit temperature, prior scale $\sigma$, and threshold $\kappa$ for DBTM, and churn and prior scale for FMLM) [10].

LM1B generative perplexity versus sample entropy frontiers for FMLM, DBTM and DBTM with ril at 1, 2, 4, 8 and 16 NFE.
Figure 22Generative frontiers on LM1B for NFE ∈ {1, 2, 4, 8, 16}. Lower Gen-PPL and higher entropy are better, and the dotted line marks the data entropy (4.336). DBTM and DBTM + ril dominate FMLM across entropies within the range of the data.
OpenWebText generative perplexity versus sample entropy frontiers for FMLM, FMLM+, DBTM and DBTM with ril at 1, 2, 4, 8 and 16 NFE.
Figure 23Generative frontiers on OpenWebText for NFE ∈ {1, 2, 4, 8, 16}, comparing FMLM, FMLM+, DBTM, and DBTM + ril. Lower Gen-PPL and higher entropy are better, and the dotted line marks the entropy of real OWT text. Each curve pools every inference-time knob for that method.

5.3Reasoning with DBTM

DBTM outperforms discrete and continuous methods on Sudoku and GSM8K in only a few NFEs

Sudoku. We measure one-shot exact-solve accuracy on 2,000 held-out puzzles at three difficulty levels. On the hard split, DBTM + ril solves 84.6% of puzzles at 4 NFE, 13.4 points above FMLM+ (71.2%). At 16 NFE it reaches 97.5%, compared with 81.4% for FMLM+. The best many-step baseline, Duo, reaches 58.4% using 128 NFE.

The trace below is a real 4-NFE sample from the paper. This illustrates the refinement loop: the first proposal is already mostly right, and each round commits the confident cells and fixes the rest against the growing context.

Figure 24A real Sudoku-Hard trace at 4 NFE. Mint cells are correct and high-confidence, so they are committed. Sky cells are correct but low-confidence, so they are renoised. Lavender cells are incorrect, so they are renoised. The number of committed cells grows from 0 to 30, 47, 64, and 81. Correctness coloring uses the known solution for explanation only, and the sampler never sees it.

GSM8K. We additionally benchmark on the harder math reasoning task GSM8K. Trained on TinyGSM code [11] and evaluated zero-shot on the 1,319-problem GSM8K test set, DBTM + ril reaches 16.8% at 32 NFE, compared with 16.6% for FMLM+. This is competitive or superior to other continuous and discrete diffusion baselines with only 32 NFEs (versus 1,024 used by the discrete and continuous diffusion baselines). However, autoregressive models remain dominant, at 53.9% (sampled) and 63.3% (greedy) with 512 NFE.

Figure 25A real TinyGSM trace (paper appendix). Rounds 1, 10, and 20 of 31. Every intermediate iterate is a complete program, so the trace can be read as it refines.
Explore Table 2Interactive 7
DBTM + rilFMLM+ (few-step flow map)many-step baselines

Sudoku bars show exact-solve accuracy (%) on 2,000 held-out puzzles, and GSM8K bars show accuracy (%) on the 1,319-problem test set. The NFE used by each method is printed inside its bar. FMLM+ hard and GSM8K numbers are from Agarwal et al. [4], and FMLM+ easy and medium at 16 NFE are our own evaluation of the released checkpoints.

5.4Comparison to flow maps for language modeling

Flow map language models and DBTM both achieve few-step generation with key differences

Flow map language models [3, 12], and DBTM share the same goal of few-step generation, but they differ in several key components, which we summarize below.

Flow map LMs / DFMDBTM (ours)
Few-step generationYesYes
Learned objectTwo-time map $X_{s,t}$Stationary endpoint map $T$
Time conditioningYes ($s$ and $t$)No
Teacher modelYesNo
Each NFE is…A discretized step along the ODEA refinement of a clean proposal
Self-stoppingNo, the step schedule is fixed in advanceYes, it stops at the fixed point or when all tokens pass $\kappa$
Training complexityHigh, with diagonal/off-diagonal balancing over all $(s,t)$Low, with one model conditioned only on the current state
Train–inference mismatchYes, many $(s,t)$ jumps are unused at inferenceLow, since interpolants cover inference states and ril training adds rollouts
Reasoning traceNoYes, every iterate is a clean sequence
Part 06

Conclusions

  1. An autonomous flow on the simplex gives straighter, time-averaged paths whose only stable equilibria are the vertices.
  2. The transport map solves a conservation equation, so it can be trained end-to-end from data, with no teacher and no time conditioning.
  3. Every application proposes a clean sequence, so extra NFEs become renoise-and-refine steps that self-correct against committed context.
  4. Across language and reasoning, DBTM improves on discrete diffusion and flow baselines at 1–16 NFE.

We’re excited to see the future extensions of this framework across modalities! For more details, read the DBTM paper as well as the BTM paper.

References

  1. Tang, Wang. Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning. arXiv 2026.
  2. Lee et al. Beckmann Transport Models: From Autonomous Flows to One-Step Maps. NeurIPS 2026.
  3. Lee et al. Flow Map Language Models: One-step Language Modeling via Continuous Denoising. NeurIPS 2026.
  4. Agarwal et al. Posterior Refinement: Fast Language Generation via Any-Order Flow Maps. arXiv 2026.
  5. Albergo et al. Stochastic Interpolants: A Unifying Framework for Flows and Diffusions. JMLR 2025.
  6. Beckmann. A Continuous Model of Transportation. Econometrica 1952.
  7. Derrida. Random-Energy Model: Limit of a Family of Disordered Models. Phys. Rev. Lett. 1980.
  8. Boffi et al. How to Build a Consistency Model: Learning Flow Maps via Self-Distillation. NeurIPS 2025.
  9. Radford et al. Language Models are Unsupervised Multitask Learners (GPT-2). OpenAI 2019.
  10. Pynadath, Shi, Zhang. Generative Frontiers: Why Evaluation Matters for Diffusion Language Models. arXiv 2026.
  11. Liu et al. TinyGSM: Achieving >80% on GSM8K with Small Language Models. arXiv 2023.
  12. Potaptchik et al. Discrete Flow Maps. arXiv 2026.

BibTeX

@article{tang2026discrete,
  title={Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning},
  author={Sophia Tang and Shiyi Wang},
  journal={arXiv preprint arXiv:2609.15903},
  year={2026},
  url={https://arxiv.org/abs/2609.15903}
}

@inproceedings{cheukkit2026beckmann,
  title={Beckmann Transport Models: From Autonomous Flows to One-Step Maps},
  author={Lee Cheuk-Kit and Florentin Coeurdoux and Yuyuan Chen and Sophia Tang and Peter Potaptchik and Yilun Du and Michael S. Albergo and Eric Vanden-Eijnden},
  booktitle={Advances in Neural Information Processing Systems},
  year={2026},
  url={https://arxiv.org/abs/2608.01692}
}