knowledge-base/raw/xpost/2026-08-19_junsong-dflash2-speculative-decoding.md
Hector 1d9c67a8e3 ingest(xpost): Qwen3.8-27B DFlash 2 speculative decoding + uncensored debate
- raw (NEW): junsong-dflash2-speculative-decoding (70 tok/s M5 Max, 4.6x, Z Lab -> Inco AI)
- raw (NEW): gregpr07-qwen38-uncensored (no gates) + s1gmoid Gegenposition
- wiki (NEW): concepts/llm/speculative-decoding.md
- wiki-update(concepts/llm): qwen3.8-27b-alibaba — DFlash 2 / Speculative Decoding + Uncensored-Debatte
- wiki/index + wiki/log aktualisiert
2026-08-19 10:01:00 +02:00

939 B
Raw Blame History

type source_url retrieved author is_thread
xpost https://x.com/jun_song/status/2089849870651892183 2026-08-19 @jun_song false

Jun Song: Qwen3.8-27B Speculative Decoding (DFlash 2)

Quelle: https://x.com/jun_song/status/2089849870651892183

Bezug: https://x.com/zhijianliu_/status/2089836737132650504 (Zhijian Liu, DFlash-2-Ankündigung)

Inhalt

Jun Song: "Qwen3.8-27b hitting 70 tok/s on a single MacBook. That is faster than Fable or Sol. Speculative decoding is easily the biggest breakthrough in local AI this year. Next up, the real innovation is going to happen in prefill and weight compression."

Bezug: Zhijian Liu (DFlash 2)

"DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro. Up to 4.6× the speed of autoregressive decoding, with the same output. This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free!"