Tokens & Signals · Friday, August 21, 2026

NVIDIA AVO Aces ARC-AGI: The 100% Benchmark Barrier Falls

ox-alphagpt-5.6-solgrok-4.6fable-5-maxavoopenai-codexclaude-mythos-5qwen-3.8-27bglm-4v-maxclaude-opus-4.6-20251101-thinking-32kgemini-3.1-proopenrouternvidiaopenaianthropicstarcloudkagicoding-agentsarc-agimodel-benchmarkingon-device-aiorbital-datacentersenterprise-securityquantizationmodel-distillationapi-pricingautonomous-agentskarpathyteortaxestexchetanp
Tokens & Signals for 8/21/2026. We scanned ~1,200 Twitter accounts (1244 tweets), 13 subreddits (62 posts), Hacker News (7 stories), 4 newsletter posts, 3 podcast episodes, 94 Discord messages, and leaderboard data for you. Estimated reading time saved: ~10 hours.

TLDR & AI Twitter Recap

* A mysterious new model called "Ox Alpha" just showed up on OpenRouter with a 1M token context window and is absolutely stomping DeepSWE at 80%—well ahead of GPT-5.6 Sol (52%). x.com/teortaxesTex/status/2090631974209626297

* @karpathy on the new 100% ARC-AGI benchmark: "A 100% score on any hard task is either cheating, leaking, or genuinely amazing. Time will tell which."

* Grok 4.6 hit #1 on CursorBench 3.2 (70.8%) and costs a laughable $2.81/task compared to Fable 5 Max's $17.32. x.com/XFreeze/status/2090839305585377458

* NVIDIA's "AVO" agent just scored a perfect 100% on ARC-AGI-3, solving all 183 levels with zero pre-programmed rules. x.com/kimmonismus/status/2090814903133098211

* OpenAI Codex hit 20M active users and is rolling out "banked resets" so devs can carry over unused API quota. x.com/thsottiaux/status/2090766694897619318

* Anthropic launched Claude Mythos 5 for enterprise security scanning and dropped $35M for open-source security work. x.com/claudeai/status/2090852314319880425

* OpenAI cut GPT-5.6 Sol API prices by over 20% for the next 3 months to get more agentic apps off the ground. x.com/OpenAIDevs/status/2090888116014137718

* Qwen 3.8-27B is flying on consumer hardware—users are hitting 60-63 tokens/s with FP4 quantization. reddit.com/r/LocalLLaMA/comments/1vub9od/fastes...

* @teortaxesTex on Ox Alpha: "Either this is GLM-4V-Max with a new name, or someone leaked something really impressive. The 1M context alone is wild." x.com/teortaxesTex/status/2090631974209626297

* Starcloud raised at a $2.3B valuation to build datacenters in orbit. Space-based AI is officially a thing. x.com/chetanp/status/2090863679701139910

* Kagi added a setting to filter out paywalled search results, finally killing the "click-then-hit-a-wall" loop. news.ycombinator.com/item?id=49388154


Go deeper on what matters to you

Tap to expand

Best to Build With Today

* Codingclaude-opus-4.6-20251101-thinking-32k holds the top spot on Chatbot Arena Coding. For IDE tasks, grok-4.6 leads CursorBench 3.2.

* Reasoninggemini-3.1-pro is the current leader for complex math and logic according to Chatbot Arena.

* Chatgemini-3.1-pro takes the #1 spot for everyday general chat.

* Open-sourceQwen 3.8-27B is the go-to for running high-performance code on consumer hardware.

* Value pickGPT-5.6 Sol API is now 20% cheaper, making it the most cost-effective option for scaling agentic workflows.


Deeper Dives

🧠 Models & Research

Stealth model 'Ox Alpha' outperforms frontier models on DeepSWE

A mysterious model called "Ox Alpha" appeared on OpenRouter with a 1M token context window and promptly made itself at home at the top of the leaderboard. It scored 80% on DeepSWE—a coding-focused benchmark—which puts it well ahead of GPT-5.6-Sol (52%) and Fable (65%).

Why it matters: Its sudden, high-performance arrival on coding benchmarks is shaking up the hierarchy of top-tier models.

� Twitter� Reddit

NVIDIA AVO agent achieves 100% on ARC-AGI-3

NVIDIA's AVO agent solved all 183 levels of ARC-AGI-3 without a single pre-programmed goal. It uses a neuro-symbolic approach that separates visual reasoning from standard language prediction, letting it crack unseen, grid-based logic puzzles cold.

Why it matters: A perfect score on ARC-AGI is a massive milestone for truly autonomous, goal-directed AI.

� Twitter� Reddit

🚀 Products & Launches

Grok 4.6 takes #1 spot on CursorBench 3.2

Grok 4.6 topped the CursorBench 3.2 leaderboard with a 70.8% score, proving it's a legitimate powerhouse for IDE-integrated coding. At just $2.81/task, the pricing isn't even close to the competition.

Why it matters: High performance at an aggressive price is a direct threat to every expensive coding model in the room.

� Twitter

Claude Mythos 5 security scans launch in enterprise beta

Anthropic is making a move into the "AI defender" space with Claude Mythos 5, which scans for threats like prompt injection in real time. It includes a transparency dashboard so teams can see exactly why the model flags something.

Why it matters: This is a clear signal that Anthropic is going after enterprise security workflows.

� Twitter

💼 Industry & Business

OpenAI reports 20M active Codex users, offers banked resets

OpenAI confirmed Codex hit 20 million active users and is introducing a "banked reset" policy—unused API quota now rolls over across billing cycles instead of disappearing into the void.

Why it matters: The adoption numbers are massive, and the quota flexibility is a real quality-of-life win for developers building high-demand apps.

� Twitter

GPT-5.6 Sol API prices reduced by over 20%

OpenAI is cutting GPT-5.6 Sol API prices by over 20% for the next three months, covering both input and output tokens. The savings come from improvements in model distillation.

Why it matters: The race to the bottom on price is heating up as everyone fights for dominance in the agentic application layer.

� Twitter

Starcloud secures $2.3B valuation for orbital datacenter mission

Starcloud just hit a $2.3B valuation with a plan to launch satellite clusters that deliver high-performance compute in low-earth orbit for latency-sensitive AI workloads.

Why it matters: The sheer amount of capital behind this shows how far companies are willing to go to solve the AI infrastructure crunch.

� Twitter


Funding & Deals

* Starcloud: Secured a $2.3B valuation to develop orbital datacenters for high-performance AI compute.

* Anthropic: Announced a $35M fund for open-source security research alongside the Claude Mythos 5 release.


Launches

* Grok 4.6: Now available via Vertex AI, sitting at the top of the CursorBench leaderboard with superior coding efficiency.

* Claude Mythos 5: Enterprise beta model featuring real-time security scanning and prompt injection defense.


Closing thought: Between orbital datacenters and agents solving logic puzzles with 100% accuracy, the "frontier" of AI is shifting from just being smart to being physically and architecturally transformative.