Tokens & Signals · Wednesday, August 12, 2026

Grok 4.6 and the Race to the Bottom

grok-4.6qwen3.8-2.4tdeepseek-v4-progpt-5.2-codexclaude-opus-4-8-xhigh-effortclaude-sonnet-5-xhigh-effortgemini-3.1-prominimax-h3grok-4.7rtx-pro-6000ltx-2.5gpt-4ospacexaialibabadeepseekgoogleopenaihugging-facetailscalelovablenvidiaagentic-modelsopen-weightschain-of-thoughtmodel-pricingsecurity-vulnerabilityfundingvideo-generationhardware-supplylegislationinference-latencykarpathyswyxsergey-brinelon-musksam-altman
Tokens & Signals for 8/12/2026. We scanned ~1,200 Twitter accounts (1244 tweets), 13 subreddits (73 posts), Hacker News (15 stories), 7 newsletter posts, 4 podcast episodes, 121 Discord messages, and leaderboard data for you. Estimated reading time saved: ~12 hours.

TLDR & AI Twitter Recap

* SpaceXAI dropped Grok 4.6, a massive agentic model hitting frontier-level performance at a fraction of the cost ($2/M input, $6/M output) thanks to a new Mixture-of-Depths architecture. x.com/MTSlive/status/2087608197444173933

* Alibaba released Qwen3.8-2.4T, one of the largest open-weights models ever built — 1M token context window, 95B active parameters. x.com/ClementDelangue/status/2087562019788697818

* DeepSeek V4 Pro came out swinging with $0.435/M input and $0.87/M output, deliberately undercutting everyone else in the room. x.com/thdxr/status/2087610161636471289

* @karpathy on the model race: "We're watching three concurrent arms races: capability, price, and speed to deployment. Most people only see the first one."

* Google's top brass, including Sergey Brin, are forcing an "all-in" pivot to accelerate Gemini development and close the gap on frontier leaders. x.com/MTSlive/status/2087601938389250397

* Researchers found they can "steal" hidden reasoning traces from black-box APIs with ~80% accuracy — turns out hidden Chain-of-Thought is a security liability, not a moat. x.com/scaling01/status/2087312337283768582

* @swyx on the reasoning extraction news: "This is the 'side-channel attack' of the LLM era—your model's Chain-of-Thought is now a liability." x.com/swyx/status/2087437017840046156

* AI startup Lovable just locked in a massive $400M Series C at a $13.3B valuation. x.com/gustaf/status/2087565549790498946

* NVIDIA's RTX Pro 6000 (Blackwell) jumped to $16k, which is a pretty blunt reminder that hardware supply is still the industry's biggest chokepoint. reddit.com/r/LocalLLaMA/comments/1vm5e14/rtx_60...

* 31 members of Congress are demanding answers from OpenAI about the "Hugging Face incident." x.com/MTSlive/status/2087318391992537214

* Tailscale traced a production bug back to a 16-year-old SQLite issue — a humbling reminder that even the boring foundational stuff has skeletons in the closet. news.ycombinator.com/item?id=49272832

Go deeper on what matters to you

Tap to expand

Best to Build With Today

* Codinggpt-5.2-codex for general logic; claude-opus-4-8-xhigh-effort for complex agentic workflows.

* Reasoningclaude-sonnet-5-xhigh-effort (LiveBench leader) or gemini-3.1-pro for Arena-style math.

* Chatgemini-3.1-pro is currently running away with the ELO leaderboards.

* Video generationMiniMax H3 is the new community favorite for prompt adherence and 3D accuracy.

* Open-sourceQwen3.8-2.4T is the undisputed heavy hitter for local/open-weights users.

* Value pickDeepSeek V4 Pro for high-throughput tasks on a budget.

Deeper Dives

🧠 Models & Research

SpaceXAI Releases Grok 4.6

Grok 4.6 uses a Mixture-of-Depths (MoD) architecture that allocates compute dynamically per token, cutting inference latency by 15% vs Grok 4.5. It hits 92.4 on MMLU and leads GPT-4o in news classification.

Why it matters: xAI is matching frontier labs while pricing like a massive disruptor.

� Twitter� Reddit

Qwen3.8-2.4T Open-Weights Released

Alibaba's dense 2.4T parameter model (95B active) ships with a 1M token context window and hits 88.5% on HumanEval.

Why it matters: It gives the open-source community a genuine high-performance alternative to closed models.

� Twitter� Reddit� Hacker News

Research: Reasoning Trace Extraction

New research shows hidden Chain-of-Thought reasoning can be reconstructed with ~80% accuracy by fine-tuning models on API logit distributions.

Why it matters: Proprietary "hidden" reasoning isn't nearly as secret as companies want to believe.

� Twitter� Reddit

Grok 4.7 Teased

Elon Musk confirmed Grok 4.7 drops in 3–4 weeks with "Reflection-Tune" training to reduce hallucinations and a 20% boost in complex logic tasks.

� Twitter

💼 Industry & Business

DeepSeek V4 Pro Price War

DeepSeek launched at $0.435/M input and $0.87/M output, officially breaking the $1/M output threshold.

Why it matters: This forces every major API provider to justify their markups out loud.

� Twitter� Reddit

Google's All-In Gemini Pivot

Sergey Brin has ordered an "all-in" pivot, merging the Gemini frontier team with Google Research to slash development cycles.

Why it matters: Google is essentially admitting its previous pace wasn't cutting it.

� Twitter

Congress vs. OpenAI

31 members of Congress sent a formal letter to Sam Altman demanding logs about the "Hugging Face incident," signaling a new era of legislative oversight.

� Twitter

🚀 Products & Launches

Lovable's $13.3B Valuation

Lovable raised $400M at a $13.3B valuation to pour capital into GPU infrastructure and R&D talent.

� Twitter� Hacker News

MiniMax H3 Video

MiniMax H3 is taking over the Stable Diffusion community for its superior prompt adherence and camera movement compared to LTX 2.5.

� Reddit

Tailscale's SQLite Bug

Tailscale traced their recent database corruption to a 16-year-old SQLite bug — a good reminder that "battle-tested" and "bug-free" are not the same thing.

� Hacker News

Funding & Deals

* Lovable raised $400M in a Series C to scale its AI-powered coding platform.

Launches

* Grok 4.6 — High-performance agentic model from xAI.

* Qwen3.8-2.4T — Alibaba's massive open-weights model.

* DeepSeek V4 Pro — The new, ultra-cheap reasoning benchmark.

Closing thought: Between sub-dollar API pricing and the ability to reconstruct "hidden" reasoning traces from the outside, the incumbents are getting squeezed from every direction. We're leaving the era of black boxes behind — and moving into a hyper-competitive market where your model's inner monologue might just be someone else's side-channel attack.