Tokens & Signals for 7/14/2026. We scanned ~1,200 Twitter accounts (1378 tweets), 13 subreddits (58 posts), Hacker News (13 stories), 5 newsletter posts, 4 podcast episodes, 92 Discord messages, and leaderboard data for you. Estimated reading time saved: ~11 hours.
* Demis Hassabis wants a US-led international body to set mandatory safety standards for frontier AI — think CERN, but for models that could end civilization. x.com/demishassabis/status/2076957440109625718
* OpenAI's GPT-5.6 Sol just cracked a 50-year-old Erdős math problem that had been stumping human researchers since the '70s. Not a small deal. reddit.com/r/singularity/comments/1uvrtl0/anoth...
* Prism-ML dropped Bonsai-27B, a mobile-ready model that runs locally on under 5GB of RAM while still keeping up with full-precision peers. x.com/adrgrondin/status/2077115680718025138
* Anthropic told the US Senate that Alibaba allegedly used 25,000 fake accounts to scrape 28.8 million private Claude conversations — apparently to clone the model's behavior. reddit.com/r/ChatGPT/comments/1uwavzo/anthropic...
* OpenAI folded Codex into the main ChatGPT Work app, resetting usage limits and cleaning up the experience for 8 million users. x.com/sama/status/2077033807736459713
* An AI agent apparently spent 48 hours improving itself with zero human help, boosting its own success rate by 22%. Recursive self-improvement is no longer just a thought experiment. x.com/zhengyaojiang/status/2077079778793042425
* Palo Alto Networks CEO Nikesh Arora wants a 90% cut in token costs, arguing that current pricing makes long-term enterprise AI adoption basically impossible to budget for. reddit.com/r/singularity/comments/1uwa1mv/well_...
* Microsoft dropped TypeScript 7.0 with a native port that's up to 10x faster. Developers are happy. x.com/code/status/2076781322929139914
* @theallinpod on model security: "We are in a race for 'greatness' that ignores safety, leading to inevitable, massive data leaks and systemic earthquake events." x.com/theallinpod/status/2077114660172808525
Best to Build With Today
* Coding — claude-opus-4-6-20251101-thinking-32k is still the top pick for complex software architecture and serious logic work.
* Reasoning — gemini-3.1-pro is leading the pack for high-level math and general reasoning right now.
* Chat — gemini-3.1-pro is your best bet for everyday assistance and conversation.
* Open-source — Bonsai-27B by Prism-ML is the new local champion for edge-optimized, mobile-friendly deployments.
Deeper Dives
🧠 Models & Research
GPT-5.6 Sol model performance
GPT-5.6 Sol just solved a 50-year-old Erdős math problem, and the method is worth paying attention to. It uses a "Self-Reflective Chain-of-Thought" (SR-CoT) approach that lets it catch its own logical errors in real-time — which is why it's outperforming predecessors on the MATH-500 benchmark.
Why it matters: Models are moving past pattern matching into something that looks a lot more like genuine mathematical research.
� Twitter� Reddit
Recursive self-improvement breakthrough
Researchers documented an autonomous agent that spent eight days training itself and ended up beating benchmarks previously held by hand-tuned systems. In just 48 hours of autonomous operation, it improved its own performance by 22% using a self-correction loop — no humans needed.
Why it matters: AI bootstrapping its own capabilities without human intervention isn't theoretical anymore. It's happening.
� Twitter
Demis Hassabis proposes Frontier AI safety body
Hassabis is pushing for a US-led global organization to oversee frontier AI, modeled loosely on CERN. The idea centers on centralized safety benchmarking and mandatory compute governance to head off catastrophic risks as we get closer to AGI.
Why it matters: There's a growing consensus that safety testing needs real structure — not just voluntary commitments from labs with obvious conflicts of interest.
� Twitter� Reddit� Hacker News
💼 Industry & Business
Anthropic reports massive scraping by Alibaba
Anthropic told the US Senate that Alibaba used 25,000 fake accounts to scrape 28.8 million Claude conversations over six weeks. The goal, Anthropic alleges, was to copy model behavior and train competing systems on proprietary dialogue data.
Why it matters: This is a serious escalation in the data war, and it raises hard questions about IP protection and cross-border data theft that nobody has good answers to yet.
� Reddit
Rising AI model costs
Enterprise AI adoption is running into a pricing wall. Palo Alto Networks CEO Nikesh Arora is pushing for a 90% price cut, saying current token costs make it basically impossible to budget for AI-integrated security at scale.
Why it matters: Inference costs are still the single biggest bottleneck to wide-scale industrial AI adoption. Until that changes, a lot of promising use cases stay stuck in pilot purgatory.
� Twitter� Reddit
Cerebras CEO warns of AI "earthquake" events
Andrew Feldman argued that frontier AI companies are prioritizing "greatness" over safety, and predicted it leads somewhere bad — think massive data leaks and systemic failures.
Why it matters: This isn't a safety researcher crying wolf. It's someone who understands the hardware infrastructure driving all of this, and he's genuinely concerned.
� Twitter
🚀 Products & Launches
OpenAI merges Codex and ChatGPT Work
OpenAI combined its developer-focused Codex app and ChatGPT Work into a single platform to cut down on context switching. Enterprise users now get shared limits and search across both.
Why it matters: Developers who need coding tools and LLM chat in the same workflow no longer have to juggle two separate products.
� Twitter� Hacker News
OpenAI encrypts sub-agent prompts
OpenAI started encrypting sub-agent prompts in Codex, which makes it harder for developers to audit what's actually happening inside agentic reasoning traces.
Why it matters: Black-boxing agentic logic is a real problem — it makes debugging harder and security auditing significantly more difficult at exactly the moment when people are trusting these systems with more.
� Hacker News
Funding & Deals
* Anthropic committed $10 million CAD to fund AI research partnerships with Canadian institutions, focused on safe and ethical AI development. anthropic.com/news/canadian-ai-research
Launches
* Bonsai-27B — Prism-ML's new 27B-parameter model built to run locally on mobile devices with under 5GB of VRAM. Efficient without sacrificing too much performance. prismml.com/news/bonsai-27b
* TypeScript 7.0 — Microsoft's performance-focused release featuring a native port that delivers up to 10x faster execution. x.com/code/status/2076781322929139914
Closing thought: Whether it's agents quietly improving themselves overnight or a 27B model squeezing into 5GB of memory, the gap between "lab demo" and "running on your phone" is shrinking fast. The capability side of AI is moving at a pace that's genuinely hard to track. The safety and infrastructure sides? Still catching up — and today's stories make that pretty clear.