Automated Daily Intelligence · Est. 2026

TIDOX.HUB
EPISODE 135 · 2026.08.04

Intelligence Brief

Swiftlet streams routed Mixture-of-Experts weights, AirLLM retook #1 on GitHub Trending with a differe, A same-day voice/TTS cluster -- livekit/agents , j

3B

parameters per token regardles

INTELLIGENCE BRIEF
August 4, 2026
DAILY EDITION
2026-08-04INTELLIGENCE BRIEF
0:00 / 5:49
Daily intelligence brief · Two-voice podcast with visuals
· RSS · Get daily email (coming soon)

Today's Insights

Swiftlet (HN Show #1) streams routed Mixture-of-Experts weights from disk, activating ~3B parameters per token regardless of total model size: Qwen3.6-35B-A3B runs at 2.6GB peak RAM (7-11 tok/s on an M5 Mac), and Qwen3-Next-80B-A3B runs at 4.3GB peak RAM (4.5-5 tok/s on an M5 Mac, ~1 tok/s at 2.5GB RAM on an iPhone 17). The developers' own caveat: these models 'chat and write like large models but recall facts like small ones.'

local-model-inferencon-devicequantization GitHub / Hacker News Show

Trending Repos