Est. MMXXVI No. 14 ◆ Archive ⌁ Free Models
LOCAL AI NEWS

Inference self-hosted on a DGX Spark · models served via LiteLLM

‘Killmonger Locs’ Have Become Gaming’s Black Hair Default, and One Creator Is Pushing Back — wired.com
An AI Agent Hacked a Home Network to Prove a Point — wired.com
MS NOW Bets on Fandom, Not Just Cable, With New Paid Membership — wired.com
‘Killmonger Locs’ Have Become Gaming’s Black Hair Default, and One Creator Is Pushing Back — wired.com
An AI Agent Hacked a Home Network to Prove a Point — wired.com
MS NOW Bets on Fandom, Not Just Cable, With New Paid Membership — wired.com

Major runtime release

vLLM 0.29.0 makes Model Runner V2 the default for all models

The latest vLLM release turns Model Runner V2 into the default everywhere and adds support for Hy4-preview, Qwen3.8-Flash-Next, GraniteSWA/MoeSWA and NemotronH_Omni_Reasoning_V3. It also ships major Kimi K3 and DeepSeek V4 kernel optimizations plus expanded speculative decoding (DSpark, MTP, sparse MLA). For anyone serving open-weight models locally, this is the most consequential runtime update of the day.

Research & Papers

05

Online draft co-training for speculative decoding in long-context RL post-training

Open Weights & Models

03

IBM releases SOTA Granite Time Series PatchTST-FM-r2 under a permissive license

Tools & Runtimes

05

Qwen ZG: local-first semantic grep for AI agents

Self-Hosted Stack

04

n8n ships 2.38.6 stable with memory-bound source-control push

Hardware & Edge

04

A modder brings NVIDIA's unreleased RTX 3070 Ti 16GB to life

Free Models

24

Free means $0 per token today, not unlimited or unconditional: both platforms rate-limit free use and require a free API key, some OpenRouter free endpoints may train on your prompts unless you opt out in account settings, and the roster changes daily. Verified against each provider's own pricing data at generation time — but pricing changes without notice, so confirm it's still free on the provider's own page before you rely on it.

In Brief

12

Source: