Tokens & Signals · Friday, August 28, 2026

Anthropic’s Agents: Outperforming Humans at Scale

glm-5.3hy4-previewgemini-omni-1.1-flashclaude-opus-5-maxclaude-fable-5gpt-image-2minimax-h3kimi-k3-maxglm-5.3-maxminimax-h3-maxanthropica16ztencentgooglegoogle-deepmindz.aiminimaxplaidgoogle-researchcoding-agentsroboticsedge-hardwareai-safetylong-contextmultimodalityopen-sourcefundingdmcareasoningkarpathy
Tokens & Signals for 8/28/2026. We scanned ~1,200 Twitter accounts (1168 tweets), 13 subreddits (71 posts), Hacker News (7 stories), 4 newsletter posts, 3 podcast episodes, 144 Discord messages, and leaderboard data for you. Estimated reading time saved: ~10 hours.

TLDR & AI Twitter Recap

* Anthropic's automated researchers are killing it: Their agents outperformed human experts at patching jailbreaks (94% vs 78%) — and wrapped it up in a single day. x.com/AnthropicAI/status/2093386528668172373(ht...)

* A16z is betting big on the physical world: They just closed a $1.1B "Machine Age Fund" to back robotics, chips, and edge hardware. x.com/MTSlive/status/2093350777804992811(https:...)

* GLM-5.3 is the new open-weights heavyweight: A 5.3T MoE model with a 50% jump in coding performance and some seriously impressive benchmark numbers. huggingface.co/zai-org/GLM-5.3(https://huggingf...)

* Tencent's Hy4-preview is here: A 1M-token context window with a sparse attention trick keeping latency in check. huggingface.co/tencent/Hy4-preview(https://hugg...)

* Pentagon blacklisting ruled illegal: A federal judge shut down the effort to block Anthropic from government contracts, calling it a straight-up due process violation. news.ycombinator.com/item?id=49473522(https://n...)

* @karpathy on AI safety: "The irony of using AI to align AI is not lost on me, but the alternative is using humans, and we have a 94% vs 78% problem now."

* Grok adds real-time finance: Track crypto and stocks via Plaid integration right inside the chat. x.com/XFreeze/status/2093337680452968873(https:...)

DMCA abuse is getting scary: The game Luanti* got yanked from the Play Store by what looks like a malicious, AI-generated copyright bot. news.ycombinator.com/item?id=49475079(https://n...)

* Gemini Omni 1.1 Flash is faster: Out now with 20% lower costs and tighter latency for voice and video apps. x.com/GoogleDeepMind/status/2093338200580256172...)

* Claude Desktop gets a CLI bridge: You can finally use /resume to pull terminal sessions straight into your desktop app. x.com/ClaudeDevs/status/2093368017304371503(htt...)


Go deeper on what matters to you

Tap to expand

Best to Build With Today

* Coding: claude-opus-5-max remains the top pick for complex web dev tasks.

* Reasoning: claude-opus-5-max is still leading the pack for high-level logic and math.

* Chat: claude-fable-5 is the crowd favorite for everyday assistant work.

* Image generation: gpt-image-2 (medium) holds the crown for quality and editing precision.

* Video generation: gemini-omni-1.1-flash leads on text-to-video, while minimax-h3 is your go-to for image-to-video.

* Open-source: glm-5.3-max and kimi-k3-max are the best bets if you want high-performance open-weights models.


Deeper Dives

🧠 Models & Research

GLM-5.3 Open-Weights Released

Z.ai's GLM-5.3 runs on a 5.3 trillion parameter MoE architecture and delivers a 15% jump in HumanEval performance. The secret sauce is a novel synthetic reasoning chain curriculum — basically proof that smart post-training can keep the closed-model labs looking over their shoulders.

Why it matters: It gives devs a genuinely top-tier model they can actually run and host themselves.

� Twitter� Reddit� Hacker News

Anthropic's Automated Alignment Breakthrough

Anthropic's Claude-powered systems can now identify and patch failure modes — like jailbreaks — better than human red-teamers, hitting a 94% success rate versus 78% for humans.

Why it matters: If AI can police its own safety boundaries, it fundamentally changes how fast AI development can move.

� Twitter

Tencent's Hy4 Preview

Tencent's 770B parameter model is built for long-context work, packing a 1M-token window and hitting 99.8% recall on needle-in-a-haystack tests thanks to a new sparse attention mechanism that keeps things fast.

Why it matters: Long-context agentic workflows are the next big frontier for coding and research, and this is a serious entry.

� Twitter� Reddit

Google's WikiSkill Framework

Google Research's WikiSkill lets agents store experiences in a persistent "wiki" instead of losing everything when a session ends. It boosted task completion by 25% on ToolBench for unseen APIs.

Why it matters: Real-world agents need long-term memory — this gives them a structured way to actually have it.

� Twitter

🚀 Products & Launches

Gemini Omni 1.1 Flash

Google's latest update is all about speed: 20% lower token costs and sharper video generation performance, which makes it a natural fit for any app where latency actually matters.

Why it matters: It meaningfully lowers the bar for building real-time voice and video apps.

� Twitter

MiniMax H3 Max

Optimized for the Fal platform, this model is now 14x faster on a single GPU — solid proof you don't need a massive render farm to get high-quality video output.

Why it matters: High-end generative video is now within reach for indie devs.

� Reddit

💼 Industry & Business

Pentagon Blacklisting Ruling

A federal judge ruled that the Pentagon's blacklisting of Anthropic was illegal — the government couldn't back up its claims with actual evidence. Anthropic is back in the running for defense contracts.

Why it matters: It sets a significant legal precedent for how the government can (and can't) target AI companies under the banner of "national security."

� Hacker News� Reddit

A16z Machine Age Fund

Andreessen Horowitz raised $1.1B to chase "Physical AI" — the hardware, chips, and robotics that represent the real-world bottleneck for the AI boom.

Why it matters: The money is shifting from pure software plays to the infrastructure that actually puts AI into the physical world.

� Twitter�️ Podcast

🔥 Takes & Drama

Luanti's DMCA Nightmare

Open-source game Luanti got pulled from Google Play by an automated DMCA claim that the team says is completely baseless.

Why it matters: It's a textbook example of how weaponized, AI-driven copyright enforcement is becoming a real threat to the open-source community.

� Hacker News� Reddit


Funding & Deals

* Andreessen Horowitz raised $1.1B for the "Machine Age Fund" to target robotics, semiconductors, and edge-computing infrastructure.


Launches

* GLM-5.3 — Massive new 5.3T parameter MoE model focused on high-end reasoning.

* Hy4 Preview — Tencent's 1M-token context model for complex research and coding.

* Gemini Omni 1.1 Flash — Optimized multimodal model for production-grade low latency.


Closing thought: With the focus shifting from raw model scale to agentic efficiency and physical hardware, the "AI summer" is finally getting its hands dirty.