--- created: 2026-06-29 updated: 2026-06-29 sources: - xpost/2026-06-29_paulcouvert-local-llm-laptop-guide.md - concepts/llm/mlx-moe-local-ai-optimization.md - concepts/hardware/cloud-exit-and-local-superiority.md - concepts/llm/llm-model-catalog.md - concepts/llm/chinese-model-cost-routing.md tags: [concept, llm, local-ai, laptop, quantization, unsloth, qwen, gemma, parakeet, llama-cpp, lm-studio, edge-ai, voice-to-text, diffusion, moe] --- # Local LLM Laptop Guide — Best Models Without a $10k Mac Studio > **TL;DR:** Paul Couvert (@itsPaulAi) lists the best local models for laptop-class hardware. Top picks: Qwen3.6-27B for coding, Gemma 4 12B for everyday tasks, Parakeet 0.6B v3 for voice-to-text, Gemma 4 E4B for phone-class edge inference, and Gemma 4 26B diffusion for maximum tok/s. Key tips: use Unsloth quantized versions, run via LM Studio or llama.cpp, try Google AI Edge Gallery on mobile. Complements Jun Song's MLX MoE optimization post — both converge on the Qwen3.6-35B-A3B MoE variant. ## Source | Source | Author | Date | Engagement | |--------|--------|------|------------| | [X-Post: "Best local models on a laptop"](https://x.com/itsPaulAi/status/2071717503105900993) | Paul Couvert (@itsPaulAi) | 2026-06-29 | 138 likes, 12 reposts, 170 bookmarks, ~13.5K views | **Shared by:** Pit Weber in OME-Gruppe, Topic "Tips & Tricks" (Topic 27) --- ## The Model Stack ### 1. Coding: Qwen3.6-27B (or Qwen3.6-35B-A3B for speed) | Dimension | Qwen3.6-27B (Dense) | Qwen3.6-35B-A3B (MoE) | |-----------|---------------------|------------------------| | **Total params** | 27B | 35B | | **Active params/token** | 27B (all active) | 3B (sparse) | | **Best for** | Coding agents — "by far the best" | Coding, but faster | | **Laptop fit** | Needs quantization (Unsloth) | MoE sparse activation = faster on limited hardware | | **Cross-ref** | Jun Song warns: slow on Mac/MLX | Jun Song's MoE preference aligns | > **⚠️ Tension with Jun Song:** Jun Song explicitly says dense 27B on Mac is "painfully slow." Paul recommends it for coding on any laptop — likely assuming Unsloth quantization on non-Mac hardware. The **MoE variant (35B-A3B)** is where both agree: it's the faster, better-suited option for constrained hardware. See [resolution below](#resolving-the-dense-27b-tension). ### 2. Everyday Tasks: Gemma 4 12B | Dimension | Detail | |-----------|--------| | **Size** | 12B | | **Best for** | Everyday tasks and questions | | **Why** | Balanced size-to-quality ratio for general use | | **Runtime** | LM Studio, llama.cpp, or Google AI Edge Gallery | ### 3. Voice-to-Text: Parakeet 0.6B v3 | Dimension | Detail | |-----------|--------| | **Size** | 0.6B — tiny | | **Best for** | Speech-to-text (STT) | | **Claim** | "Basically the best voice-to-text model you can find" | | **Relevance** | Local Whisper alternative; relevant for OpenClaw voice interaction pipelines | | **Edge fit** | Small enough to run on any laptop or phone | ### 4. Edge / Phone: Gemma 4 E4B | Dimension | Detail | |-----------|--------| | **Size** | ~4B | | **Best for** | On-device inference, offline phone use | | **Claim** | "Surprisingly good for the size" | | **Mobile deployment** | Google AI Edge Gallery (iOS/Android) | | **Cross-ref** | Extends [Cloud-Exit thesis](../hardware/cloud-exit-and-local-superiority.md) to mobile — connects to [OpenClaw mobile apps](../../tools/openclaw.md) | ### 5. Speed Champion: Gemma 4 26B Diffusion | Dimension | Detail | |-----------|--------| | **Size** | 26B | | **Type** | Diffusion-based LLM | | **Best for** | Maximum tok/s on local hardware | | **Claim** | "The most tok/s you can have with a local model" | | **Architecture note** | Diffusion LLMs generate tokens via iterative denoising — different paradigm from standard AR models. See [[post-transformer-llm-architectures.md]] for DeepMind's diffusion LLM research (10× efficiency gains). | --- ## Practical Deployment Tips ### Quantization: Unsloth [Unsloth](https://github.com/unslothai/unsloth) provides quantized model versions optimized for laptop deployment: | Benefit | Detail | |---------|--------| | **Precision/size/speed** | "Fantastic" balance — Paul's explicit recommendation | | **Compatibility** | Works with LM Studio, llama.cpp | | **How to use** | Download Unsloth quantized GGUF files → load in LM Studio or llama.cpp | ### Runtime Options | Runtime | Type | Best For | |---------|------|----------| | [LM Studio](https://lmstudio.ai) | GUI app | Easy setup, model browsing, chat interface | | [llama.cpp](https://github.com/ggerganov/llama.cpp) | CLI / library | Maximum control, minimal overhead, embedding in apps | | [Google AI Edge Gallery](https://ai.google.dev/edge) | Mobile app | On-device inference on iOS/Android | --- ## Resolving the Dense 27B Tension Paul recommends Qwen3.6-27B (dense) for coding. Jun Song says dense 27B on Mac is "painfully slow." This isn't a contradiction — it's a **context distinction**: | Factor | Paul's Context | Jun Song's Context | |--------|---------------|-------------------| | **Hardware** | Any laptop (x86, ARM, Mac) | Specifically Apple Silicon / MLX | | **Quantization** | Unsloth quantized (explicitly recommended) | Dequantized (dq) MoE models | | **Framework** | llama.cpp / LM Studio (CPU/GPU offload) | MLX (Apple native) | | **Dense 27B verdict** | ✅ Good with Unsloth quants | ❌ "Painfully slow" on MLX | | **MoE 35B-A3B verdict** | ✅ "Faster option" | ✅ Implied by MoE preference | **Resolution:** The performance of dense 27B depends heavily on **framework + quantization**: - **Unsloth quantized on llama.cpp** (Paul's path): tractable on commodity laptops - **Unquantized/dq on MLX** (Jun Song's path): slow because MLX doesn't optimize for dense - **MoE variant (35B-A3B)**: optimal on both paths — sparse activation reduces compute per token **Recommendation:** If you're on a Mac, follow Jun Song's advice (MoE models via MLX). If you're on any laptop with llama.cpp + Unsloth quants, Paul's dense 27B recommendation works. When in doubt, pick the MoE variant (Qwen3.6-35B-A3B) — it's the convergence point. --- ## Complement to MLX MoE Optimization This page is the **cross-platform companion** to [[mlx-moe-local-ai-optimization.md]]: | Dimension | [MLX MoE Optimization](mlx-moe-local-ai-optimization.md) | This Page (Laptop Guide) | |-----------|----------------------------------------------------------|--------------------------| | **Scope** | Apple Silicon / MLX | Any laptop (x86, ARM, Mac) | | **Key insight** | MoE > Dense on MLX (unified memory) | Quantized models work on any laptop | | **Model sources** | Minimax-M3.0, DeepSeek-v4-Flash (Chinese MoE) | Qwen3.6, Gemma 4, Parakeet (broader mix) | | **Focus** | Why MoE works on Apple Silicon | What to run + how (runtime, quantization) | | **Unique coverage** | Info gap thesis (frontier AIs can't tell you) | STT (Parakeet), phone-edge (Gemma E4B), diffusion speed (Gemma 26B) | **Together:** Jun Song tells you *why* MoE on Mac is better. Paul tells you *what* to run on any laptop and *how*. The MoE variant Qwen3.6-35B-A3B is the overlap — both recommend it, for different reasons. --- ## Model Landscape Summary | Use Case | Model | Size | Architecture | Hardware Floor | |----------|-------|------|--------------|----------------| | Coding (best) | Qwen3.6-27B | 27B | Dense | Laptop w/ quantization | | Coding (fast) | Qwen3.6-35B-A3B | 35B/3B active | MoE | Laptop | | Everyday | Gemma 4 12B | 12B | Dense | Laptop | | Voice-to-text | Parakeet 0.6B v3 | 0.6B | STT | Any device | | Phone/edge | Gemma 4 E4B | ~4B | Dense | Phone | | Speed champion | Gemma 4 26B diffusion | 26B | Diffusion LLM | Laptop (GPU helpful) | --- ## Connection to Our Setup ### Model Routing Alignment | Model | Paul's Role | Our Role | Match | |-------|-------------|----------|-------| | Qwen3.6-35B-A3B | "Faster coding option" (MoE) | Not yet in active routing — candidate | Potential addition | | Qwen3.6-27B | "Best for coding agents" | Not in active routing | Reference point | | Gemma 4 12B | "Everyday tasks" | Not in active routing | Consider for lightweight tasks | ### Cloud-Exit Extension Paul's guide extends the Cloud-Exit thesis to **laptop-class hardware** — you don't need a Mac Studio, DGX Station, or Strix Halo: | Hardware Tier | Cloud-Exit Option | Source | |---------------|-------------------|--------| | Phone | Gemma 4 E4B via Google AI Edge Gallery | Paul Couvert | | Laptop (any) | Qwen3.6/Gemma 4 with Unsloth quants via LM Studio/llama.cpp | Paul Couvert | | Mac (Apple Silicon) | MoE models (Minimax-M3.0, DeepSeek-v4-Flash) via MLX | Jun Song | | Prosumer desktop | AMD Strix Halo (128 GB) | [[../hardware/edge-inference-als-cloud-alternative.md]] | | Enterprise desktop | NVIDIA DGX Station (748 GB) | [[../hardware/nvidia-dgx-station-748gb.md]] | --- ## Cross-References - [[mlx-moe-local-ai-optimization.md]] — Jun Song's MLX MoE post (direct complement, Apple Silicon focus) - [[../hardware/cloud-exit-and-local-superiority.md]] — Cloud-Exit thesis (this page extends to laptop tier) - [[../hardware/edge-inference-als-cloud-alternative.md]] — AMD Strix Halo (edge inference, x86) - [[../hardware/nvidia-dgx-station-748gb.md]] — NVIDIA DGX Station (enterprise tier) - [[llm-model-catalog.md]] — Consolidated model catalog - [[chinese-model-cost-routing.md]] — DeRonin's cost routing (Qwen models in cloud→local swap) - [[post-transformer-llm-architectures.md]] — Diffusion LLMs (Gemma 4 26B diffusion context) - [[../../tools/openclaw.md]] — OpenClaw mobile apps (Gemma E4B on-phone connection) - [[ai-search-language-blind-spot.md]] — "Local" als Querschnittsthema: Quellen-Routing (lokale Sprache) analog zu Edge-Inferenz (lokale Hardware) ## External Links - [Original X-Post by Paul Couvert](https://x.com/itsPaulAi/status/2071717503105900993) - [Unsloth GitHub](https://github.com/unslothai/unsloth) - [LM Studio](https://lmstudio.ai) - [llama.cpp](https://github.com/ggerganov/llama.cpp) - [Google AI Edge Gallery](https://ai.google.dev/edge)