Est. MMXXVI No. 6 ◆ Archive ⌁ Free Models
LOCAL AI NEWS

Inference self-hosted on a DGX Spark · models served via LiteLLM

Nvidia reintroduces DLSS 5 neural rendering after a March backlash — theverge.com
Nvidia’s DLSS 5 launches this week with heavy GPU demands — theverge.com
Running a Chatbot on Your Own Computer — wired.com
Nvidia CEO Took Trump’s Call Mid-All-Hands, Sources Say — wired.com
Nvidia reintroduces DLSS 5 neural rendering after a March backlash — theverge.com
Nvidia’s DLSS 5 launches this week with heavy GPU demands — theverge.com
Running a Chatbot on Your Own Computer — wired.com
Nvidia CEO Took Trump’s Call Mid-All-Hands, Sources Say — wired.com

Spark chips, deepened

NVIDIA and MediaTek pledge multiple generations of RTX Spark and DGX Spark PC chips

NVIDIA and MediaTek are expanding their partnership to build several future generations of RTX Spark and DGX Spark PC chips, pairing NVIDIA GPUs (including Vera Rubin and upcoming Feynman architectures) with MediaTek SoCs for consumer PCs and workstations. MediaTek also adopts NVIDIA's NVLink Fusion for custom AI infrastructure. A multi-generation commitment to affordable, local AI hardware.

Research & Papers

05

A universal context-reuse layer for cross-model KV sharing

Open Weights & Models

02

NVIDIA quantizes Qwen3.8-2.4T-A95B to 4-bit NVFP4 — roughly 2.5x smaller on vLLM

Tools & Runtimes

05

shaide: K8s-native AI platform for distributed multi-model inference

Free Models

24

Free means $0 per token today, not unlimited or unconditional: both platforms rate-limit free use and require a free API key, some OpenRouter free endpoints may train on your prompts unless you opt out in account settings, and the roster changes daily. Verified against each provider's own pricing data at generation time — but pricing changes without notice, so confirm it's still free on the provider's own page before you rely on it.

In Brief

07

Source: