Est. MMXXVI No. 18 ◆ Archive ⌁ Free Models
LOCAL AI NEWS

Inference self-hosted on a DGX Spark · models served via LiteLLM

OpenAI Skipping Its IPO in 2026, Altman Confirms — theverge.com
Y Combinator's Garry Tan Backs Distillation Approach for US Open-Weight AI Labs — techcrunch.com
Wired's Best Mushroom Coffees of 2026: A Coconut Latte Takes the Top Spot — wired.com
Blender Plug-in Speeds Up Creation of Textured Black Hair in Games — wired.com
OpenAI Skipping Its IPO in 2026, Altman Confirms — theverge.com
Y Combinator's Garry Tan Backs Distillation Approach for US Open-Weight AI Labs — techcrunch.com
Wired's Best Mushroom Coffees of 2026: A Coconut Latte Takes the Top Spot — wired.com
Blender Plug-in Speeds Up Creation of Textured Black Hair in Games — wired.com

Memory squeeze

The AI boom is now a line item on your next laptop receipt

Memory price trackers show the HBM4 ramp starving commodity DRAM: Samsung and SK Hynix fell below 10 days of finished DRAM inventory on September 7, and HBM3E 36GB 12-Hi stacks listed at $300.87 (+14% year on year) on September 13. A 32GB DDR5 kit that cost $80-120 in early 2025 was near $529 by mid-2026. New fab output from Samsung, SK Hynix and Micron is not expected to bring meaningful relief before mid-2027 or 2028, so plan a local rig budget around sustained memory inflation.

vLLM AOT config pushes Qwen3.8-27B INT4 to a 147K-token window on one 24GB card

A measured recipe posts Qwen3.8-27B (INT4 AutoRound) on vLLM 0.27.1 with an FP8 E4M3 KV cache at a 147,456-token window on a single RTX 3090, hitting about 872 tok/s prefill and 38 tok/s decode against roughly 25-30 tok/s from llama.cpp with MTP. The author notes AOT compilation resolves Inductor/Triton first-startup OOMs, and that GDN linear-attention state pools make vLLM's VRAM estimates unreliable.

Open Weights & Models

04

DeepSeek V4 Flash Vision Exp arrives as open weights

Tools & Runtimes

03

Kairo routes LLM inference with a fail-closed layer calibrated on RTX 5090 measurements

Hardware & Edge

04

High-end AMD gaming bundle lands under $1,110 for a complete local-inference desktop

Free Models

25

Free means $0 per token today, not unlimited or unconditional: both platforms rate-limit free use and require a free API key, some OpenRouter free endpoints may train on your prompts unless you opt out in account settings, and the roster changes daily. Verified against each provider's own pricing data at generation time — but pricing changes without notice, so confirm it's still free on the provider's own page before you rely on it.

In Brief

07

Source: