TL;DR
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Q Labs Research describes Dust, a zeroth-order method that trains transformers by perturbing activations rather than calculating gradients through backpropagation. The authors report competitive results in their experiments and better computational efficiency than a weight-space evolution-strategy method, while saying Dust requires larger populations and more compute than backprop in some comparisons. Independent validation and evidence at larger scales are not established in the supplied report.
Q Labs Research has introduced Dust, a method for pretraining transformer language models without backpropagation, reporting that it can compete with backprop in experiments by searching over model activations. The work matters because backpropagation is the standard way to train large neural networks; the authors say Dust offers an alternative, but its reported performance depends on population size and additional computation, and the findings have not been independently established in the supplied material.
Dust is a zeroth-order optimization method: instead of computing a model’s gradient through a backward pass, it perturbs activations and uses the resulting changes in loss to estimate how parameters should be updated. The report says perturbations are applied independently at each token. That lets the method treat tokens as members of a virtual population and evaluate them together during a forward pass, rather than constructing and separately running many perturbed copies of the model.
Q Labs says Dust’s estimates become more aligned with backpropagation as the population grows and remain well aligned across the scales tested, up to 1 billion tokens. The report also says a 243 million-parameter model outperformed a model 120 times smaller at most tested population sizes. These are results reported by the authors; the supplied material does not establish that they have been independently reproduced.
The authors compare Dust with EGGROLL, a weight-space evolution-strategy method. They estimate that from 1 million tokens onward Dust is roughly 1,000 to 10,000 times more efficient in their comparison. That figure is an extrapolation, not a direct demonstration of a matching advantage across all model sizes or training conditions. The report also acknowledges that Dust approximates backprop closely at large populations, which it describes as requiring substantially more compute.
A Different Route to Transformer Training
If the results hold up, Dust could broaden the methods researchers can use to train neural networks, particularly where compute is abundant but reliance on differentiability or conventional gradient calculations is a constraint. The authors argue that activation-space search can make large populations practical by evaluating perturbations together, rather than paying the cost of separately materializing every candidate network.
The report does not show that Dust is ready to replace backprop in routine model training. Its claims concern research experiments and include comparisons whose cost depends on population size and the evaluation setup. For developers and researchers, the near-term significance is a proposed alternative worth testing—not evidence that current training pipelines should change.
high performance GPU for AI training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Weight Search to Activations
Most modern transformer training uses backpropagation, which calculates how changes to model parameters affect a loss function. Evolution strategies take a different approach: they perturb weights, evaluate the resulting models, and use those outcomes to guide updates. Q Labs characterizes such methods as costly at scale because each population member must be represented and evaluated.
Dust draws on node perturbation, applying changes to activations instead of weights. Its proposed virtual population relies on the many tokens processed by a transformer, with each token’s perturbation contributing information during a shared forward pass. The report’s central claim is that this arrangement can reduce the cost of population-based search. The supplied source is a Q Labs Research report dated October 2026; it does not provide independent assessment of the method.
“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”
— Q Labs Research, in the report’s abstract
As an affiliate, we earn on qualifying purchases.
Questions Beyond the Reported Tests
The supplied report does not establish independent replication, performance on a broad range of language-model benchmarks, or results at the scale of major production pretraining runs. It also does not settle how Dust’s total compute and hardware costs compare with backprop when both methods are trained to the same target quality under matched conditions.
The efficiency comparison with EGGROLL is described as an extrapolation, and the report’s claim of competitiveness depends in part on population size. It remains unclear how sensitive results are to data, model architecture, tuning choices, and training duration, or whether the method maintains its reported behavior beyond the tested scale.
As an affiliate, we earn on qualifying purchases.
Replication and Larger-Scale Comparisons
The next useful evidence would be independent reproductions of Dust’s training results and transparent, matched-compute comparisons with backprop and evolution-strategy baselines. Such tests could clarify whether the reported advantage is specific to the authors’ setup or persists across models, datasets, and hardware.
The report points toward scaling tests, but the supplied material does not announce a release schedule, external evaluation, or a follow-up milestone. Until those details are available, Dust should be treated as a research proposal with promising author-reported results rather than a confirmed replacement for standard transformer training.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Dust?
Dust is a zeroth-order training method from Q Labs Research. It perturbs transformer activations and uses changes in loss to estimate updates without backpropagation.
Does Dust eliminate the need for backpropagation?
The report describes training without a backward pass, but it does not establish that Dust can replace backpropagation in general-purpose or production-scale training. Its competitiveness is an author-reported experimental result.
How does Dust use tokens as a population?
According to Q Labs, Dust perturbs activations independently at each token and treats each token as a virtual population member. The authors say one forward pass evaluates these members in parallel.
What does the report say about efficiency?
Q Labs estimates that Dust is about 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from 1 million tokens onward. The report labels this an extrapolation, and the comparison is not a universal measure of Dust’s cost against backpropagation.
What remains unproven?
The supplied material does not show independent replication, matched-compute results across a wide range of models, or performance at large production-training scales. Those tests would help establish whether the reported findings generalize.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
