The project GitHub page is a much better introduction for the hn crowd.
HoldOnAMinute 5 minutes ago [-]
How is this different from other LLM runners?
simoiacos 2 hours ago [-]
Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.
I'm also looking into expanding the protocol and the engine to support various steering techniques.
It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
pulkitsh1234 8 minutes ago [-]
curious, why did antirez go with C instead of something like Rust ?
GTP 5 minutes ago [-]
Personal preference of the author, he made at least one video on YouTube on why he dislikes Rust. I think he finds it too cumbersome and not worth it when the software isn't security-critical (not that I agree, just reporting what IIRC his stance is).
doctorpangloss 3 hours ago [-]
the problem is the dsv4 checkpoint so quantized isn't very good
ilaksh 2 minutes ago [-]
Which ds4 checkpoint for which model exactly did you test? Don't they have multiple different versions and quantization levels?
The project GitHub page is a much better introduction for the hn crowd.
I'm also looking into expanding the protocol and the engine to support various steering techniques.
https://github.com/simoneiacomino/xenolith
the ds4 quants were very good beating the unsloth quants https://github.com/michaelasper/benchmarks/blob/main/deepsee...
What are we going to name the company, how about Dwarfism 2.0? What happened to 1.0 Jared?