91 lines
4.8 KiB
Markdown
91 lines
4.8 KiB
Markdown
|
|
---
|
||
|
|
type: xpost
|
||
|
|
source_url: https://x.com/itsPaulAi/status/2071717503105900993
|
||
|
|
author: "@itsPaulAi"
|
||
|
|
date: 2026-06-29
|
||
|
|
posted_by: Pit Weber (@PWeber)
|
||
|
|
topic: "Tips & Tricks (Topic 27)"
|
||
|
|
tags: [xpost, local-llm, laptop, qwen, gemma, parakeet, unsloth, llama-cpp, lm-studio, quantization, voice-to-text, diffusion, edge-ai, google-ai-edge-gallery]
|
||
|
|
---
|
||
|
|
|
||
|
|
# Paul Couvert: Best Local Models You Can Run on a Laptop Today
|
||
|
|
|
||
|
|
**Source:** [X-Post by @itsPaulAi](https://x.com/itsPaulAi/status/2071717503105900993)
|
||
|
|
**Posted:** 2026-06-29
|
||
|
|
**Shared by:** Pit Weber in OME-Gruppe, Topic "Tips & Tricks" (Topic 27)
|
||
|
|
|
||
|
|
## Post Content
|
||
|
|
|
||
|
|
Paul Couvert (@itsPaulAi) lists the best local models for laptop-class hardware — explicitly noting you don't need a $10k Mac Studio:
|
||
|
|
|
||
|
|
> The best local models you can run on a laptop today (without a $10k Mac Studio):
|
||
|
|
>
|
||
|
|
> 1. **Qwen3.6-27B** — by far the best for coding agents (or **Qwen3.6-35B-A3B** for a faster option)
|
||
|
|
> 2. **Gemma 4 12B** — everyday tasks and questions
|
||
|
|
> 3. **Parakeet 0.6B v3** — basically the best voice-to-text model you can find
|
||
|
|
> 4. **Gemma 4 E4B** — surprisingly good for the size, can run offline on a phone
|
||
|
|
> 5. **Gemma 4 26B diffusion** — the most tok/s you can have with a local model
|
||
|
|
>
|
||
|
|
> Tips:
|
||
|
|
> - Use **quantized versions from Unsloth** when running on a laptop — precision/size/speed are fantastic
|
||
|
|
> - **LM Studio** or raw **llama.cpp** are the best runtime options
|
||
|
|
> - **Google AI Edge Gallery** on iOS/Android
|
||
|
|
|
||
|
|
## Model Overview Table
|
||
|
|
|
||
|
|
| # | Model | Size | Use Case | Key Strength |
|
||
|
|
|---|-------|------|----------|--------------|
|
||
|
|
| 1 | Qwen3.6-27B | 27B | Coding agents | "By far the best for coding" |
|
||
|
|
| 1b | Qwen3.6-35B-A3B | 35B (MoE, 3B active) | Coding (faster alt) | MoE sparse activation for speed |
|
||
|
|
| 2 | Gemma 4 12B | 12B | Everyday tasks/questions | General-purpose balanced |
|
||
|
|
| 3 | Parakeet 0.6B v3 | 0.6B | Voice-to-text (STT) | "Best voice-to-text model you can find" |
|
||
|
|
| 4 | Gemma 4 E4B | ~4B | Edge / phone | Runs offline on phone |
|
||
|
|
| 5 | Gemma 4 26B diffusion | 26B | Speed-focused | "Most tok/s for a local model" |
|
||
|
|
|
||
|
|
## Practical Tips
|
||
|
|
|
||
|
|
| Tip | Detail |
|
||
|
|
|-----|--------|
|
||
|
|
| **Quantization** | Use Unsloth quantized versions — "precision/size/speed are fantastic" |
|
||
|
|
| **Runtime** | LM Studio or raw llama.cpp recommended |
|
||
|
|
| **Mobile** | Google AI Edge Gallery app (iOS/Android) for on-device inference |
|
||
|
|
|
||
|
|
## Engagement
|
||
|
|
|
||
|
|
- 138 likes
|
||
|
|
- 12 reposts
|
||
|
|
- 18 replies
|
||
|
|
- 170 bookmarks
|
||
|
|
- ~13.5K views
|
||
|
|
|
||
|
|
## Context: Complements Jun Song MLX MoE Post
|
||
|
|
|
||
|
|
This post complements the [Jun Song MLX MoE post](https://x.com/jun_song/status/2071625328083227118) (wikified as [[concepts/llm/mlx-moe-local-ai-optimization.md]]):
|
||
|
|
|
||
|
|
| Dimension | Jun Song | Paul Couvert |
|
||
|
|
|-----------|----------|-------------|
|
||
|
|
| **Focus** | Apple Silicon / MLX optimization | Any laptop (cross-platform) |
|
||
|
|
| **Key insight** | MoE > Dense on MLX | Quantized models from Unsloth for any laptop |
|
||
|
|
| **Top coding pick** | Minimax-M3.0 (dq), DeepSeek-v4-Flash (dq) | Qwen3.6-27B (or 35B-A3B MoE variant) |
|
||
|
|
| **Dense 27B on Mac** | "Painfully slow" — avoid | Recommends Qwen3.6-27B for coding (on any laptop, quantized) |
|
||
|
|
| **Resolution** | Different deployment contexts: Jun Song = Mac/MLX native; Paul = any laptop with Unsloth quants | |
|
||
|
|
| **Common ground** | Both recommend Qwen3.6-35B-A3B (MoE variant) — Jun Song implicitly (MoE preference), Paul explicitly (faster option) | |
|
||
|
|
|
||
|
|
**Key distinction:** Jun Song warns about dense 27B on Mac/MLX specifically. Paul recommends Qwen3.6-27B for coding agents on any laptop — likely with Unsloth quantization, which changes the performance profile. The MoE variant (Qwen3.6-35B-A3B) is where both converge: Jun Song's MoE preference + Paul's "faster option."
|
||
|
|
|
||
|
|
## Relevance to Our Setup
|
||
|
|
|
||
|
|
1. **Qwen3.6-35B-A3B MoE variant** — aligns with Jun Song's MoE recommendation and our model routing (MoE models as fallbacks)
|
||
|
|
2. **Unsloth quantization** — practical tip for running models on commodity hardware, complements Cloud-Exit thesis
|
||
|
|
3. **Parakeet 0.6B for STT** — relevant for OpenClaw voice interaction, potential local Whisper alternative
|
||
|
|
4. **Gemma 4 E4B on phone** — edge inference extends Cloud-Exit to mobile, connects to OpenClaw mobile apps
|
||
|
|
5. **LM Studio / llama.cpp** — runtime options already mentioned in our wiki (GLM-5.2 self-hosting via LM Studio)
|
||
|
|
|
||
|
|
## Related Wiki Pages
|
||
|
|
|
||
|
|
- `wiki/concepts/llm/mlx-moe-local-ai-optimization.md` — Jun Song's MLX MoE post (direct complement)
|
||
|
|
- `wiki/concepts/hardware/cloud-exit-and-local-superiority.md` — Cloud-Exit thesis
|
||
|
|
- `wiki/concepts/llm/llm-model-catalog.md` — Consolidated model catalog
|
||
|
|
- `wiki/concepts/llm/chinese-model-cost-routing.md` — DeRonin's cost routing (same Qwen models)
|
||
|
|
- `wiki/concepts/hardware/edge-inference-als-cloud-alternative.md` — Edge inference (AMD Strix Halo)
|
||
|
|
- `wiki/tools/hermes-desktop.md` — Hermes MoA (related ensemble pattern)
|