Est. MMXXVI No. 8 ◆ Archive ⌁ Free Models
LOCAL AI NEWS

Inference self-hosted on a DGX Spark · models served via LiteLLM

Meta Pushes Its New AI Agent on Employees—but Eases Off on Tokenmaxxing — wired.com
TechCrunch expands Disrupt 2026 with a new Real World AI Stage — techcrunch.com
Democrats Plan to Leverage Funding to Force Trump Associates to Cooperate With Investigations — wired.com
Google's new Gemini 3.8 Flash 'works harder' but might cost more — theverge.com
US government files brief backing OpenAI in NYT copyright lawsuit — techcrunch.com
Meta Pushes Its New AI Agent on Employees—but Eases Off on Tokenmaxxing — wired.com
TechCrunch expands Disrupt 2026 with a new Real World AI Stage — techcrunch.com
Democrats Plan to Leverage Funding to Force Trump Associates to Cooperate With Investigations — wired.com
Google's new Gemini 3.8 Flash 'works harder' but might cost more — theverge.com
US government files brief backing OpenAI in NYT copyright lawsuit — techcrunch.com

Local inference speedup

Unsloth v0.1.806-beta: 2x faster Qwen3.8-Flash and GLM-5.3-Flash via MTP

Unsloth's new beta runs Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with multi-token prediction enabled by default, bundled with 170+ improvements across training, chat, hardware support, and APIs in its local inference and fine-tuning studio. It is the most directly actionable speedup this week for anyone serving the leading open-weight models on their own rig.

Research & Papers

05

Language Models Can Control Their Own Attention

Free Models

25

Free means $0 per token today, not unlimited or unconditional: both platforms rate-limit free use and require a free API key, some OpenRouter free endpoints may train on your prompts unless you opt out in account settings, and the roster changes daily. Verified against each provider's own pricing data at generation time — but pricing changes without notice, so confirm it's still free on the provider's own page before you rely on it.

In Brief

09

Source: