Tokens & Signals for 7/15/2026. We scanned ~1,200 Twitter accounts (1248 tweets), 13 subreddits (59 posts), Hacker News (11 stories), 4 newsletter posts, 2 podcast episodes, 138 Discord messages, and leaderboard data for you. Estimated reading time saved: ~10 hours.
* Thinking Machines dropped "Inkling," a 975B-parameter Mixture-of-Experts model with 41B active params, a 1M-token context window, and an Apache 2.0 license. It's their shot at giving open-source developers genuine frontier-level power. huggingface.co/thinkingmachines/Inkling-NVFP4
* OpenAI launched "Codex Micro," a $230 physical control deck built with Work Louder. It's got RGB status keys and a physical dial to crank your agent's reasoning effort up or down. openai.com/supply/co-lab/work-louder
* New York became the first US state to ban new AI data center construction in certain zones, citing energy and water concerns. A $4B project is already on ice. reddit.com/r/singularity/comments/1ux7h4t/new_y...
* @gdb on GPT-Red: "Automated, scalable red teaming is a critical piece of infrastructure for agentic safety." x.com/gdb/status/2077464463251554327
* DeepSeek is rumored to be eyeing a 2026 IPO at a $20B+ valuation. They've raised $7.4B and are already pulling in $500M in annual revenue. x.com/jukan05/status/2077248730819113410
* Claude Code now supports Model Context Protocol (MCP) connectors, letting you pipe real-time data from private databases and local tools straight into your coding workflow. x.com/ClaudeDevs/status/2077489907350856038
* @karpathy on the new Physics Olympiad results: "When AI ties top human performance on unimpeachable external benchmarks, the 'but can it really reason?' debate gets a lot quieter."
* Meta AI's Muse Spark scored a perfect 30/30 on the Asian Physics Olympiad, showing it can handle complex, multi-step scientific reasoning better than anything we've seen before. x.com/AIatMeta/status/2077138553210028042
* Local LLM users are seeing a solid 20-25% speed bump thanks to the new ExLlamaV3 v1.0.0 release.
Best to Build With Today
* Coding — claude-opus-4-6-thinking-32k (The current gold standard for complex, agentic engineering).
* Reasoning — gemini-3.1-pro (Leading the pack for math and logic problem solving).
* Chat — gemini-3.1-pro (Top-ranked on the Chatbot Arena for everyday tasks).
* Open-source — Inkling (If you need a high-performance base for domain-specific fine-tuning).
* Value pick — gemini-2.5-flash-preview (Punches way above its weight for cost-effective agentic tasks).
Deeper Dives
🧠 Models & Research
Thinking Machines Releases 'Inkling'
Inkling is a 975B-parameter Mixture-of-Experts foundation model with 41B active parameters, natively supporting text, image, and audio reasoning across a 1M-token context window. The Apache 2.0 license means companies can fine-tune it on their own private data with zero vendor lock-in.
* Why it matters: It's a massive win for the open-weights community — a real, credible alternative to closed-source frontier models.
� Twitter� Reddit
OpenAI Releases GPT-Red for Automated Security
GPT-Red is an internal adversarial model that automates security by inventing its own prompt-injection attacks against agents. In stress tests, it found 3x more jailbreaks than manual red teams in 48 hours.
* Why it matters: As agents start doing actual work for us, we can't rely on humans to hunt down every vulnerability. We need AI to guard the AI.
� Twitter� Reddit
Meta AI's 'Muse Spark' Hits 30/30 on Physics Olympiad
The model used a hybrid chain-of-thought architecture to solve the Asian Physics Olympiad theoretical exam. A perfect 30/30 is a serious milestone — it's hard to argue "pattern matching" when you're acing physics problems at that level.
* Why it matters: It's an independent, un-hackable benchmark, and the score speaks for itself.
� Twitter
Tencent's 'RxBrain' for Embodied AI
This new multimodal model bridges language reasoning with spatial imagination, designed to help AI agents understand and actually navigate physical environments.
* Why it matters: Embodied AI — models that can "do" things in the real world, not just talk about them — is shaping up to be the next big frontier.
� Reddit
German Consortium Releases 'Soofi S'
A 30B open-source model that outperforms its peers in both English and German.
* Why it matters: It's a solid reminder that specialized, multilingual open models can absolutely go toe-to-toe with the US-centric tech giants.
� Reddit� Twitter
🚀 Products & Launches
OpenAI Launches 'Codex Micro' Hardware
The $230 "Codex Micro" is a compact control deck for agentic coding, with RGB status keys and a physical dial to dial reasoning effort up or down.
* Why it matters: It's OpenAI's first move into hardware, and it signals they want to own the "vibe coding" workstation experience end-to-end.
� Twitter� Hacker News
Claude Code Adds MCP Connector Support
Claude Code can now natively pull data from local files, databases, and APIs.
* Why it matters: It turns Claude from a static generator into a dynamic system that can actually "see" and interact with your private environment.
� Twitter
💼 Industry & Business
New York Data Center Moratorium
Governor Kathy Hochul signed a one-year pause on new hyperscale data center permits, giving the state time to develop real standards around energy consumption and local infrastructure strain.
* Why it matters: It's the first major legal pushback against AI infrastructure build-out — and it'll almost certainly ripple through other states.
� Reddit
DeepSeek Eyes 2026 IPO
Reports suggest DeepSeek is preparing for a 2026 IPO with a projected $20B+ valuation.
* Why it matters: A successful public listing would cement them as a genuine global player capable of going head-to-head with US labs.
� Twitter
OpenAI Loses EU Trademark Appeal
The European Court ruled that "OpenAI" is a descriptive term for the industry rather than a distinct brand, denying them exclusive name protection in the EU.
* Why it matters: It's a real brand management headache in a market they can't afford to fumble.
� Hacker News
Launches
* ExLlamaV3 v1.0.0 — A major update to the local inference engine that boosts batch performance by 20-25% on consumer GPUs.
* Google Gemma 4 — Updated with Flash Attention 4 support and improved tool-use capabilities.
Closing thought: The industry is clearly shifting from "bigger is better" to "smarter and more controllable." Between the physical controls of the Codex Micro and the infrastructure brakes being pulled in New York, we're watching an industry finally reckon with the fact that this stuff lives in the real world — and the real world has limits.