- raw (NEW): junsong-dflash2-speculative-decoding (70 tok/s M5 Max, 4.6x, Z Lab -> Inco AI) - raw (NEW): gregpr07-qwen38-uncensored (no gates) + s1gmoid Gegenposition - wiki (NEW): concepts/llm/speculative-decoding.md - wiki-update(concepts/llm): qwen3.8-27b-alibaba — DFlash 2 / Speculative Decoding + Uncensored-Debatte - wiki/index + wiki/log aktualisiert
939 B
939 B
| type | source_url | retrieved | author | is_thread |
|---|---|---|---|---|
| xpost | https://x.com/jun_song/status/2089849870651892183 | 2026-08-19 | @jun_song | false |
Jun Song: Qwen3.8-27B Speculative Decoding (DFlash 2)
Quelle: https://x.com/jun_song/status/2089849870651892183
Bezug: https://x.com/zhijianliu_/status/2089836737132650504 (Zhijian Liu, DFlash-2-Ankündigung)
Inhalt
Jun Song: "Qwen3.8-27b hitting 70 tok/s on a single MacBook. That is faster than Fable or Sol. Speculative decoding is easily the biggest breakthrough in local AI this year. Next up, the real innovation is going to happen in prefill and weight compression."
Bezug: Zhijian Liu (DFlash 2)
"DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro. ⚡ Up to 4.6× the speed of autoregressive decoding, with the same output. This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free!"