Est. MMXXVI No. 12 ◆ Archive ⌁ Free Models
LOCAL AI NEWS

Inference self-hosted on a DGX Spark · models served via LiteLLM

ICE subpoena sought REI buyers of green beanie in Minnesota, Wired reports — wired.com
ICE subpoena sought REI buyers of green beanie in Minnesota, Wired reports — wired.com

On-device silicon

Arm puts neural accelerators inside the Mali GPU and doubles SME2 in the CPU cluster

Arm's new CSS for Mobile 2 platform pairs the Mali G2-Ultra GPU, which carries dedicated neural accelerators for the first time, with C2 CPU clusters that double the SME2 vector capability. Arm says the doubled SME2 delivers a 70% speedup on the latest small language models and that C2-Ultra reaches up to 1.7x higher AI performance than C1-Ultra while using up to 38% less power at the same speed. It is the on-device inference baseline the next generation of phones and edge boxes will be built on.

Research & Papers

02

When quantization breaks memory: recurrent-state write-back fails in low-precision temporal inference

Open Weights & Models

03

OpenBMB ships MiniCPM5-2B: the sub-4B crown goes Apache-2.0, no launch event

OpenBMB ships MiniCPM5-2B: the sub-4B crown goes Apache-2.0, no launch event

A standard LlamaForCausalLM 2B that any Llama-family engine loads today, shipping the full stack: BF16 checkpoint with -Base/-Midtrain/-SFT companions, GGUF quants down to 1.56 GB, MLX builds for Apple Silicon, and a first-party 324M DSpark speculative-decoding draft model. Artificial Analysis scores it 15 on Intelligence Index v4.2, the highest of any open-weights model under 4B parameters; benchmark figures are vendor-run, and the 512K context promised in July shipped as 131K.

IFM's K2 Horizon 36B-A4B lands in GGUF: 36B in memory, 4B active per token

The GGUF build completed 2026-09-07 makes the fully open K2 Horizon fleet locally testable tonight on a 24 GB card, with the Mixture-of-Value-Attention design activating roughly 4B of its 37.4B parameters per token and a 524K-token context window built into the architecture under Apache-2.0. Its claims of beating 30B-class dense models and MoE rivals up to 15x its size are IFM's own, not independently reproduced.

Qwen3.8-Max is open-weight, but 'open' is carrying a custom license

The 2.4T-parameter, 95B-active MoE's weights are downloadable but under a custom license, not the family's usual Apache-2.0: operators generating more than US$50M in revenue within twelve months owe Alibaba a commercial licence. Unsloth puts the cheapest usable configuration at roughly 450 GB of combined RAM and VRAM, so 'open' here functions as source transparency and API pricing leverage rather than a local-run option.

Tools & Runtimes

05

llmash: an Ollama-style local LLM interface claiming 2-4x faster decoding

Self-Hosted Stack

03

New community benchmark 'The Struggle Bench' targets local models that fail mid-task

Hardware & Edge

04

GEEKOM chains four Ryzen AI Max+ 395 mini PCs into one local inference node at IFA 2026

Free Models

22

Free means $0 per token today, not unlimited or unconditional: both platforms rate-limit free use and require a free API key, some OpenRouter free endpoints may train on your prompts unless you opt out in account settings, and the roster changes daily. Verified against each provider's own pricing data at generation time — but pricing changes without notice, so confirm it's still free on the provider's own page before you rely on it.

In Brief

12

Source: