Est. MMXXVI No. 5 ◆ Archive ⌁ Free Models
LOCAL AI NEWS

Inference self-hosted on a DGX Spark · models served via LiteLLM

The Theragun Sense makes everyday recovery surprisingly easy — techcrunch.com
Nvidia’s edge now extends beyond the chips themselves — techcrunch.com
Local LLMs can run offline on your own hardware, WIRED reports in its guide to getting started with locally hosted chatbots. — wired.com
Chinese automakers pile into humanoid robots after Tesla’s lead — techcrunch.com
Donald Trump's call cuts into Nvidia's all-hands meeting — wired.com
The Theragun Sense makes everyday recovery surprisingly easy — techcrunch.com
Nvidia’s edge now extends beyond the chips themselves — techcrunch.com
Local LLMs can run offline on your own hardware, WIRED reports in its guide to getting started with locally hosted chatbots. — wired.com
Chinese automakers pile into humanoid robots after Tesla’s lead — techcrunch.com
Donald Trump's call cuts into Nvidia's all-hands meeting — wired.com

Research & Papers

04

Layer bit-width allocation tunes LLM quantization for latency under a quality budget

Tools & Runtimes

03

llama.cpp b10715 fuses the DFlash encoder into the KV-cache decode path

Self-Hosted Stack

02

Why distributed inference hasn't taken off for 1.5TB open-weight models

Free Models

24

Free means $0 per token today, not unlimited or unconditional: both platforms rate-limit free use and require a free API key, some OpenRouter free endpoints may train on your prompts unless you opt out in account settings, and the roster changes daily. Verified against each provider's own pricing data at generation time — but pricing changes without notice, so confirm it's still free on the provider's own page before you rely on it.

In Brief

01

Source: