NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
▲Dust: Pretraining Transformers Without Backpropagation (qlabs.sh)
blt 51 minutes ago [-]
Every few years, a derivative-free neural network optimization algorithm gets some hype. I'd bet my life savings that none of them ever make an impact.

Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.

The gradient is useful. Instead of trying random directions and hoping that one of them is an improvement, it tells you where to go. The more parameters you have, the more useful it becomes.

A strict complexity gap between gradient-based and derivative-free Lipschitz convex optimization has been suspected for decades and recently proved (using AI, [2]). Neural net optimization is nonconvex, but not radically different.

IMO, a more promising direction is gradient-based optimizers specialized to the neural network structure, like Muon [3].

[1] https://arxiv.org/abs/2202.00817

[2] https://arxiv.org/abs/2607.13335

[3] https://jeremybernste.in/writing/deriving-muon

vatsachak 45 minutes ago [-]
Random selection seems to favor post training

https://arxiv.org/pdf/2603.12228

usernametaken29 2 hours ago [-]
> There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradient

Both algorithms are bound by the same Pareto frontier based on the Empirical Risk Minimisation Principle, so they’re already on the same trajectory. Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly). So removing this limitation is actually a great step. I’m excited to see a comeback of evolutionary methods because they’re much more general, albeit costly and naive. We’re now very close to what can be described best as brute forcing the Pareto frontier out of our datasets. Not sure that’s what we want but I have no better ideas either.

stevemk14ebr 2 hours ago [-]
Pareto frontier the Pareto frontier and then with a little Pareto frontier you might get to the Pareto frontier
usernametaken29 1 hours ago [-]
P^2
AIorNot 52 minutes ago [-]
Its just an insufferable AI way to say best set of options given
tomrod 36 minutes ago [-]
Waaay older than AI. Basic econ 101,it is.
usernametaken29 42 minutes ago [-]
Not really. The Pareto frontier is itself a distribution over all optimisation paths and dominance is a real and useful property we test when researching evolutionary methods. It’s not chosen to be insufferable but to try to be precise.
wg0 49 minutes ago [-]
It's computationally expensive and infeasible but what's the upside?

Genuine question due to unfamiliarity with the subject.

polyomino 6 hours ago [-]
Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory
6 hours ago [-]
6 hours ago [-]
AIorNot 48 minutes ago [-]
Hmm I was more interested in this alternative to backprop:

https://news.ycombinator.com/item?id=49701182

api 7 hours ago [-]
It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
vatsachak 7 hours ago [-]
Not necessarily, backprop is highly parallelizable since it is just a bunch of matrix mults.

Something like Dust skips the backward pass on backprop. But other techniques like Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".

The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.

Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params

tyromaniac 5 hours ago [-]
Also has the "advantage" of being slightly more biologically plausible as the optimization happens locally rather than globally.

That idea was taken further by N'dri et al in PCL, in which "activation energy" was minimized as well, and inhibitory neurons added https://www.nature.com/articles/s41467-025-64234-z.pdf

While trying to find the link for that I stumbled upon

https://arxiv.org/pdf/2605.12732

Which also looks pretty interesting

ACCount39 5 hours ago [-]
The reason why I don't see the promise for ML-only applications is that the coordination backprop requires comes very cheap to us.

"Much easier given the right device" - the "right" there just isn't shaped like the devices we actually build. And the price of "not having backprop" is usually expending more FLOPs, getting worse sample efficiency, etc.

The biggest "device" that doesn't do backprop is the brain, and that's because the brain doesn't have the connectivity or the coordination to pull it off. Both of those are "expensive" for something like it to implement. Cheap for us though. We aren't stuck with neurons that only get locally available information and have to implement learning rules based on that. So, skill issue?

vatsachak 5 hours ago [-]
I mean co-ordination requires energy though. The brain wattage looks at GPUs and says "skill issue". But you're right that we look at natural energy production techniques and say "skill issue"
ACCount39 4 hours ago [-]
Even today's LLMs suddenly get power-competitive when you compare by "power per task". Sure, a GPU can draw 1000W under load. But it also works very fast, and doesn't have to spend any time on things like "sleep".

The trick about comparing the two is that different things are expensive to different substrates.

Coordination is cheap for GPGPU and expensive for brain. When you have a fixed number of reusable general purpose computational units, coordinating execution is more natural than not coordinating execution, and the power cost is nil. When your computational units are independent, purpose specific, and fully embedded into the data path, coordinating them can get less natural and, frankly, optional. When wiring is expensive, coordination can become expensive in turn.

Another thing in the same "cheap for GPGPU but expensive for brain" regime is bandwidth. Look no further than optic nerve to see just how hard it is for nerves to push any appreciable amount of data. Another thing is connectivity. For GPGPU, global connectivity is natural - but the brain has to pay in physical wires for all the connectivity it has, and, see "bandwidth": it doesn't have any good wires. Yet another thing is weight reuse: a big part of why humans get "handedness" is that the brain can't just reuse the motion control circuitry for one hand for another nearly identical hand.

And the final thing I can name off the top of my head is memory - but specifically, memory capable of fast R/W. The capacity of human "working memory" is a disgrace, and not because there was no use for more. Humans rapidly lose visual fidelity of representations for objects they aren't directly looking at, and not at all because "being able to check how things looked 2 seconds ago" is useless. Those capabilities were just too expensive for the substrate to afford them easily.

It's why brains, broadly, favor dataflow-like and SSM-like dynamics, with largely fixed asynchronous dataflows and recurrence over updated local information - instead of something that would require a lot of global connectivity and transformer-like many-to-many attention ops. SSM is not necessarily the "best" tool for the job in ML land, for most jobs - but when you struggle to fit "attention" into your connectivity/bandwidth budget, and your memory is extremely expensive but hard-coupled to processing, SSM starts looking very appealing.

Now, something that might be expensive for GPGPU but cheap for brain, for once? Online learning. Maybe it's substrate dependent, or maybe it's going to get cheap in GPU land too once we figure out the trick. But so far? No one figured out how to make it cheap, stable and usable. You'd be lucky to get "pick one".

vatsachak 4 hours ago [-]
I mean the reason why models are using less energy is because they are getting smarter per token and also engineering algorithms/chips that make inference cheaper.

If we could have success with spiking neural networks in silico they would take even less energy, because they don't require global co-ordination. Co-ordination is information and "information = energy by the second law of thermodynamics" is my crank proof

Also the brain has way more parameters than LLMs and also has different neurotransmitters, loops, branching etc so they probably have WAY more capacity than LLMs.

But coding output/W LLMs have us beat

janalsncm 5 hours ago [-]
Knowing nothing about this, I wonder if it could be useful in situations where we can’t reliably sync with all the workers. Something like folding@home, where all the workers are just shaking weights and if one of them finds a winner it uploads to the central server?
vatsachak 5 hours ago [-]
No real advantage over Neural Nets here; backprop matmuls can be calculated layer by layer so you can chunk backprop across different machines. The real advantage comes from energy savings, you require no global co-ordination
janalsncm 4 hours ago [-]
Imagine we did that, split up a model layers as A->B->C. C will need to wait for B to compute a forward pass, which is waiting for A to compute its forward pass. To compute the forward pass, B needs all of the outputs from A, which is an upload and a download (maybe these can be done concurrently).

Then A waits for B to compute its backwards pass, which is waiting for C to do the same thing. Again you are sending around potentially gigabytes of data.

This is in contrast to mining bitcoins for example which doesn’t require any coordination from miners because their work is completely independent, and the answer is very small compared to the work needed to get it.

vatsachak 4 hours ago [-]
Yep. You need to transport all the weights at the boundary regardless of Backprop/NPC.

But the cool thing is that if your NN is split into mostly self contained chunks then you can go widthwise parallel.

An architecture like MOE exploits this fact so that the active weights during pre-training you're backproping only through active experts

spindump8930 3 hours ago [-]
The problem (and contrast with other approaches) is that mat muls requires synchronization. Arranging your networking and training structure to maximize compute and minimize communication is the main craft of ML training infra folks. In your example, yes you can compute layers on different machines (i.e. Tensor Parallelism), but you must be very careful in how you arrange it.
strbean 5 hours ago [-]
Would these alternatives to backprop make it more feasible to have constant live-training going on in a model? Giving it something akin to neuro-plasticity?
nbutton762 3 hours ago [-]
One aspect of how current training and continual learning are somewhat at odds is that the memory required to train a model is often times 2-3x the memory required to just run it (probably not as bad for PEFT, not sure).

DUST does have an advantage specifically along those lines because it doesn't have to save a ton of intermediate state other than each layer's input activations during a single forward pass.

There are many other issues that this algorithm does not address thoigh like catastrophic forgetting. it's still operating on a transformer which contains no inherent mechanism for selecting the relative value of a training step based on current knowledge, nor does it have segmentation of functionalities with specialized areas used for specific things that can be sequestered off and ignore new updates (we do not risk forgetting how to walk as we increase our French vocabulary)

vatsachak 3 hours ago [-]
Models suffer from "catastrophic forgetting" if you train them on new data.

People are working on this field, recent results suggest that continual learning can be possible by converting the input data to "LLMese"

SerdarGl 3 hours ago [-]
At massive scales 0th order methods will parallelize better than backprop especially along depth, u can train very deep models pipeline parallel without bubbles
oofbey 57 minutes ago [-]
This is super yawn-worthy. Instead of backprop for the exact gradient you can run forward passes a thousand times with perturbed weights and get a Monte Carlo estimate of the gradient. Not very clever. Extremely NOT useful.

But I guess the industry is littered with techniques for computing the same thing but vastly slower that some people find interesting. Homomorphic encryption. Zero knowledge proofs. Blockchain computing. Except in those cases there might be a legitimate reason to use it occasionally.

56 minutes ago [-]
csmlab_notes 20 minutes ago [-]
[flagged]
wrecked_em 4 hours ago [-]
Definitely more than meets the eye.
lin7c 2 hours ago [-]
[flagged]
derin-picment 7 hours ago [-]
[flagged]
eriwang915 4 hours ago [-]
Dust's 243M model beating a 120x smaller one at most population sizes is the surprising part; bigger nets got more population-efficient, not less.
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 04:56:06 GMT+0000 (Coordinated Universal Time) with Vercel.