Tokens & Signals for 8/12/2026. We scanned ~1,200 Twitter accounts (1244 tweets), 13 subreddits (73 posts), Hacker News (15 stories), 7 newsletter posts, 4 podcast episodes, 121 Discord messages, and leaderboard data for you. Estimated reading time saved: ~12 hours.
* SpaceXAI dropped Grok 4.6, a massive agentic model hitting frontier-level performance at a fraction of the cost ($2/M input, $6/M output) thanks to a new Mixture-of-Depths architecture. x.com/MTSlive/status/2087608197444173933
* Alibaba released Qwen3.8-2.4T, one of the largest open-weights models ever built — 1M token context window, 95B active parameters. x.com/ClementDelangue/status/2087562019788697818
* DeepSeek V4 Pro came out swinging with $0.435/M input and $0.87/M output, deliberately undercutting everyone else in the room. x.com/thdxr/status/2087610161636471289
* @karpathy on the model race: "We're watching three concurrent arms races: capability, price, and speed to deployment. Most people only see the first one."
* Google's top brass, including Sergey Brin, are forcing an "all-in" pivot to accelerate Gemini development and close the gap on frontier leaders. x.com/MTSlive/status/2087601938389250397
* Researchers found they can "steal" hidden reasoning traces from black-box APIs with ~80% accuracy — turns out hidden Chain-of-Thought is a security liability, not a moat. x.com/scaling01/status/2087312337283768582
* @swyx on the reasoning extraction news: "This is the 'side-channel attack' of the LLM era—your model's Chain-of-Thought is now a liability." x.com/swyx/status/2087437017840046156
* AI startup Lovable just locked in a massive $400M Series C at a $13.3B valuation. x.com/gustaf/status/2087565549790498946
* NVIDIA's RTX Pro 6000 (Blackwell) jumped to $16k, which is a pretty blunt reminder that hardware supply is still the industry's biggest chokepoint. reddit.com/r/LocalLLaMA/comments/1vm5e14/rtx_60...
* 31 members of Congress are demanding answers from OpenAI about the "Hugging Face incident." x.com/MTSlive/status/2087318391992537214
* Tailscale traced a production bug back to a 16-year-old SQLite issue — a humbling reminder that even the boring foundational stuff has skeletons in the closet. news.ycombinator.com/item?id=49272832
Best to Build With Today
* Coding — gpt-5.2-codex for general logic; claude-opus-4-8-xhigh-effort for complex agentic workflows.
* Reasoning — claude-sonnet-5-xhigh-effort (LiveBench leader) or gemini-3.1-pro for Arena-style math.
* Chat — gemini-3.1-pro is currently running away with the ELO leaderboards.
* Video generation — MiniMax H3 is the new community favorite for prompt adherence and 3D accuracy.
* Open-source — Qwen3.8-2.4T is the undisputed heavy hitter for local/open-weights users.
* Value pick — DeepSeek V4 Pro for high-throughput tasks on a budget.
Deeper Dives
🧠 Models & Research
SpaceXAI Releases Grok 4.6
Grok 4.6 uses a Mixture-of-Depths (MoD) architecture that allocates compute dynamically per token, cutting inference latency by 15% vs Grok 4.5. It hits 92.4 on MMLU and leads GPT-4o in news classification.
Why it matters: xAI is matching frontier labs while pricing like a massive disruptor.
� Twitter� Reddit
Qwen3.8-2.4T Open-Weights Released
Alibaba's dense 2.4T parameter model (95B active) ships with a 1M token context window and hits 88.5% on HumanEval.
Why it matters: It gives the open-source community a genuine high-performance alternative to closed models.
� Twitter� Reddit� Hacker News
Research: Reasoning Trace Extraction
New research shows hidden Chain-of-Thought reasoning can be reconstructed with ~80% accuracy by fine-tuning models on API logit distributions.
Why it matters: Proprietary "hidden" reasoning isn't nearly as secret as companies want to believe.
� Twitter� Reddit
Grok 4.7 Teased
Elon Musk confirmed Grok 4.7 drops in 3–4 weeks with "Reflection-Tune" training to reduce hallucinations and a 20% boost in complex logic tasks.
� Twitter
💼 Industry & Business
DeepSeek V4 Pro Price War
DeepSeek launched at $0.435/M input and $0.87/M output, officially breaking the $1/M output threshold.
Why it matters: This forces every major API provider to justify their markups out loud.
� Twitter� Reddit
Google's All-In Gemini Pivot
Sergey Brin has ordered an "all-in" pivot, merging the Gemini frontier team with Google Research to slash development cycles.
Why it matters: Google is essentially admitting its previous pace wasn't cutting it.
� Twitter
Congress vs. OpenAI
31 members of Congress sent a formal letter to Sam Altman demanding logs about the "Hugging Face incident," signaling a new era of legislative oversight.
� Twitter
🚀 Products & Launches
Lovable's $13.3B Valuation
Lovable raised $400M at a $13.3B valuation to pour capital into GPU infrastructure and R&D talent.
� Twitter� Hacker News
MiniMax H3 Video
MiniMax H3 is taking over the Stable Diffusion community for its superior prompt adherence and camera movement compared to LTX 2.5.
� Reddit
Tailscale's SQLite Bug
Tailscale traced their recent database corruption to a 16-year-old SQLite bug — a good reminder that "battle-tested" and "bug-free" are not the same thing.
� Hacker News
Funding & Deals
* Lovable raised $400M in a Series C to scale its AI-powered coding platform.
Launches
* Grok 4.6 — High-performance agentic model from xAI.
* Qwen3.8-2.4T — Alibaba's massive open-weights model.
* DeepSeek V4 Pro — The new, ultra-cheap reasoning benchmark.
Closing thought: Between sub-dollar API pricing and the ability to reconstruct "hidden" reasoning traces from the outside, the incumbents are getting squeezed from every direction. We're leaving the era of black boxes behind — and moving into a hyper-competitive market where your model's inner monologue might just be someone else's side-channel attack.