Est. MMXXVI No. 9 ◆ Archive ⌁ Free Models
LOCAL AI NEWS

Inference self-hosted on a DGX Spark · models served via LiteLLM

IFA 2026 memory

AMD launches Ryzen AI Max PRO 400 with 192GB unified memory - up to 160GB addressable as VRAM

At its IFA 2026 opening keynote AMD commercially launched the Ryzen AI Max PRO 400 series, whose single LPDDR5X-8533 pool of up to 192GB (around 273 GB/s) lets the integrated Radeon 8060S iGPU map up to 160GB as VRAM - on paper enough to hold quantized 300B-parameter models on a desk with no discrete GPU and no cloud. Commercial systems from HP and Lenovo are targeted for Q3 2026; the 192GB SKU's tight SK Hynix LPDDR5X supply remains the open question.

Research & Papers

06

Why Gated DeltaNet survives 4-bit quantization: NVFP4 W4A4 for the recurrent half of a hybrid 27B LLM

Quantizes all 496 Gated DeltaNet layers of the hybrid 27B Qwen3.8-27B to NVFP4 W4A4, including the decay and write-strength gates previously kept at 8/16-bit, and matches BF16 within seed noise across perplexity, MMLU-Pro, GSM8K, AIME'25, GPQA-Diamond and RULER to 64K tokens - disproving the community intuition that recurrent-state errors accumulate fatally under low precision.

Open Weights & Models

03

Tools & Runtimes

05

Self-Hosted Stack

02

Hardware & Edge

04

Free Models

25

Free means $0 per token today, not unlimited or unconditional: both platforms rate-limit free use and require a free API key, some OpenRouter free endpoints may train on your prompts unless you opt out in account settings, and the roster changes daily. Verified against each provider's own pricing data at generation time — but pricing changes without notice, so confirm it's still free on the provider's own page before you rely on it.

In Brief

10