local-inference 4
- My personal LLM toolchain: switch between models and harnesses, control them from your phone
- Recipe for running Qwen3.8-Flash-Next in NVFP4 on 2x RTX Pro 6000
- Running DeepSeek-V4-Flash at 700 tokens/s on 2x RTX Pro 6000
- Coding locally with Pi Coding agent and open weights models (April 2026 edition)