Tokens & Signals · Thursday, August 13, 2026

GPT-5.6 Ultrafast: 750 Tokens Per Second Is Here

gemini-3.7-flashdeepseek-v4-progpt-5.6-solclaude-opus-4-6-thinking-autoclaude-sonnet-5-xhigh-effortgemini-3-progpt-5.5-xhighqwen3.8-27bminimax-music-3googledeepseekopenaicerebrasanthropiccursorfiretigerdatabricksmetacoding-agentson-device-aiinference-speedopen-sourcecybersecuritymultimodalityobservabilityai-regulationwafer-scale-hardwarekarpathydhhjiahui-yu
Tokens & Signals for 8/13/2026. We scanned ~1,200 Twitter accounts (1324 tweets), 13 subreddits (76 posts), Hacker News (9 stories), 7 newsletter posts, 5 podcast episodes, 95 Discord messages, and leaderboard data for you. Estimated reading time saved: ~12 hours.

TLDR & AI Twitter Recap

* Google dropped Gemini 3.7 Flash, an 'intelligent workhorse' that cuts costs by 50% while absolutely dominating coding benchmarks. x.com/GoogleAIStudio/status/2087949211564183730

* DeepSeek-V4-Pro is live, and it comes with an open-source 'Harness' framework that lets agents build their own plugins on the fly. x.com/deepseek_ai/status/2087887408440164663

* OpenAI and Cerebras are testing an 'Ultrafast' mode for GPT-5.6 Sol — 750 tokens per second on wafer-scale hardware. That's not a typo. x.com/testingcatalog/status/2087951653374677152

* ChatGPT's macOS app now has 'Computer History,' which watches your screen and tracks clicks to give agents way more context about what you're actually doing. x.com/OpenAI/status/2087996496088297746

* The White House is exempting U.S.-made open-source models from mandatory safety tests to keep domestic firms competitive. x.com/kimmonismus/status/2087667694619484401

* Cursor acquired the Firetiger team to bake production-level observability directly into their AI coding agents. x.com/cursor_ai/status/2087991786279251993

* @karpathy on the new inference speed: "14x throughput is how you make agents actually useful. Latency kills agents." x.com/karpathy/status/2087951653374677152

* @dhh on the new economics: "Frontier performance at commodity pricing. The economics are flipping." x.com/dhh/status/2087867270479351885

* Anthropic is taking heat for adding invisible, persistent watermarks to all Claude outputs — and professionals who rely on anonymity are not happy about it. reddit.com/r/ClaudeAI/comments/1vn3342/claude_i...

* The White House gave vetted private firms the green light to conduct offensive cyber operations against foreign criminal groups. reddit.com/r/singularity/comments/1vn0oww/white...

Go deeper on what matters to you

Tap to expand

Best to Build With Today

* Codingclaude-opus-4-6-thinking-auto (Current leader for complex agentic tasks).

* Reasoningclaude-sonnet-5-xhigh-effort (Top performance on LiveBench).

* Chatgemini-3-pro (Highest ELO on the daily Chatbot Arena).

* Mathgpt-5.5-xhigh (Clear winner for heavy-duty analysis).

* Value pickgemini-3-7-flash (50% cheaper, high coding speed).

Deeper Dives

🚀 Products & Launches

Google Launches Gemini 3.7 Flash

Google's latest model is laser-focused on coding and agentic workflows, and comes with a 50% price cut. It scores 65.3% on DeepSWE v1.1, making it a serious option if you want high performance without burning through your budget.

* Why it matters: Google is going hard after the high-volume agent market.

� Twitter� Hacker News

OpenAI Previews 'Ultrafast' GPT-5.6 Sol Mode

Teaming up with Cerebras, OpenAI is pushing inference to 750 tokens per second — 14x faster than standard. It's in limited enterprise preview right now, with firms like Jane Street already kicking the tires.

* Why it matters: Specialized hardware is finally killing the latency bottleneck that's been holding agents back.

� Twitter� Hacker News

OpenAI Integrates 'Computer History' into Desktop App

The ChatGPT macOS app now records your app activity and clicks to build richer context for agents. It's opt-in, and it replaces the older 'Chronicle' tool.

* Why it matters: This is the shift from ChatGPT-as-chatbot to ChatGPT-as-the-thing-running-in-the-background-of-your-whole-day.

� Twitter

🧠 Models & Research

DeepSeek Releases V4 Pro and Agent Harness

The 1.6T parameter MoE model ships alongside an open-source tool called 'Harness' — a modular plugin architecture that lets agents pick up new capabilities on the fly.

* Why it matters: They're not just releasing a model; they're pushing toward self-evolving runtime environments.

� Twitter� Reddit

💼 Industry & Business

White House to Include Open Models in Safety Testing

The administration is wrapping up a cybersecurity review framework, but U.S.-developed open-weight models get a pass — a deliberate move to protect domestic competitiveness.

* Why it matters: It keeps the U.S. open-source ecosystem from getting buried under regulation.

� Twitter

White House Authorizes Private Cyberattacks

A new memorandum lets private U.S. companies run offensive cyber operations against foreign criminal groups, with federal oversight attached.

* Why it matters: The line between state-sponsored defense and private-sector digital warfare just got a lot blurrier.

� Twitter� Reddit

Anthropic Claude Watermarking Controversy

Anthropic is embedding invisible, persistent watermarks into Claude outputs, and the backlash from professionals who depend on anonymity is real and growing.

* Why it matters: Push power users too hard and they'll just go find a model that doesn't do this.

� Reddit

Cursor Acquires Firetiger

Cursor snagged the Firetiger team to wire production observability straight into their AI agents.

* Why it matters: AI coding tools are growing up — it's not just about writing code anymore, it's about keeping live systems running.

� Twitter

Funding & Deals

* Databricks raised $5 billion at a $190 billion valuation, fueled by an 80% year-over-year revenue growth.

* Jiahui Yu (Meta's multimodal lead) left to start a company focused on "a problem that will matter deeply to humanity's future."

Launches

* Qwen3.8-27B — A new open-weights model available on Hugging Face.

* MiniMax Music 3 — High-fidelity generative audio model with full ComfyUI integration.


Closing thought: The speed of iteration in agent-focused infrastructure — from wafer-scale hardware to self-evolving frameworks — makes one thing pretty clear: the industry has moved on from chatting. It's building autonomous workers now.