Est. MMXXVI No. 17 ◆ Archive ⌁ Free Models
LOCAL AI NEWS

Inference self-hosted on a DGX Spark · models served via LiteLLM

Sam Altman says OpenAI IPO in 2026 would be ill-advised — theverge.com
LG pushes back on TV data-upload claims — theverge.com
Sam Altman says OpenAI IPO in 2026 would be ill-advised — theverge.com
LG pushes back on TV data-upload claims — theverge.com
Editor's note — Light news day — fewer fresh items than usual.

Research & Papers

05

Structured transforms cut quantization overhead from O(N²) to O(N log N)

Open Weights & Models

02

Qwen3.6-based delib/sonnet fine-tune lands on Hugging Face

Tools & Runtimes

04

llama.cpp b10937 extends OpenCL row-alignment rule to more quant formats

Hardware & Edge

02

Tom's Hardware benchmarks Qwen 3.8 on RTX 5090 — VRAM alone can't beat software bottlenecks

Free Models

25

Free means $0 per token today, not unlimited or unconditional: both platforms rate-limit free use and require a free API key, some OpenRouter free endpoints may train on your prompts unless you opt out in account settings, and the roster changes daily. Verified against each provider's own pricing data at generation time — but pricing changes without notice, so confirm it's still free on the provider's own page before you rely on it.

In Brief

11

Source: