Tokens & Signals for 8/21/2026. We scanned ~1,200 Twitter accounts (1244 tweets), 13 subreddits (62 posts), Hacker News (7 stories), 4 newsletter posts, 3 podcast episodes, 94 Discord messages, and leaderboard data for you. Estimated reading time saved: ~10 hours.
* A mysterious new model called "Ox Alpha" just showed up on OpenRouter with a 1M token context window and is absolutely stomping DeepSWE at 80%—well ahead of GPT-5.6 Sol (52%). x.com/teortaxesTex/status/2090631974209626297
* @karpathy on the new 100% ARC-AGI benchmark: "A 100% score on any hard task is either cheating, leaking, or genuinely amazing. Time will tell which."
* Grok 4.6 hit #1 on CursorBench 3.2 (70.8%) and costs a laughable $2.81/task compared to Fable 5 Max's $17.32. x.com/XFreeze/status/2090839305585377458
* NVIDIA's "AVO" agent just scored a perfect 100% on ARC-AGI-3, solving all 183 levels with zero pre-programmed rules. x.com/kimmonismus/status/2090814903133098211
* OpenAI Codex hit 20M active users and is rolling out "banked resets" so devs can carry over unused API quota. x.com/thsottiaux/status/2090766694897619318
* Anthropic launched Claude Mythos 5 for enterprise security scanning and dropped $35M for open-source security work. x.com/claudeai/status/2090852314319880425
* OpenAI cut GPT-5.6 Sol API prices by over 20% for the next 3 months to get more agentic apps off the ground. x.com/OpenAIDevs/status/2090888116014137718
* Qwen 3.8-27B is flying on consumer hardware—users are hitting 60-63 tokens/s with FP4 quantization. reddit.com/r/LocalLLaMA/comments/1vub9od/fastes...
* @teortaxesTex on Ox Alpha: "Either this is GLM-4V-Max with a new name, or someone leaked something really impressive. The 1M context alone is wild." x.com/teortaxesTex/status/2090631974209626297
* Starcloud raised at a $2.3B valuation to build datacenters in orbit. Space-based AI is officially a thing. x.com/chetanp/status/2090863679701139910
* Kagi added a setting to filter out paywalled search results, finally killing the "click-then-hit-a-wall" loop. news.ycombinator.com/item?id=49388154
Best to Build With Today
* Coding — claude-opus-4.6-20251101-thinking-32k holds the top spot on Chatbot Arena Coding. For IDE tasks, grok-4.6 leads CursorBench 3.2.
* Reasoning — gemini-3.1-pro is the current leader for complex math and logic according to Chatbot Arena.
* Chat — gemini-3.1-pro takes the #1 spot for everyday general chat.
* Open-source — Qwen 3.8-27B is the go-to for running high-performance code on consumer hardware.
* Value pick — GPT-5.6 Sol API is now 20% cheaper, making it the most cost-effective option for scaling agentic workflows.
Deeper Dives
🧠 Models & Research
Stealth model 'Ox Alpha' outperforms frontier models on DeepSWE
A mysterious model called "Ox Alpha" appeared on OpenRouter with a 1M token context window and promptly made itself at home at the top of the leaderboard. It scored 80% on DeepSWE—a coding-focused benchmark—which puts it well ahead of GPT-5.6-Sol (52%) and Fable (65%).
Why it matters: Its sudden, high-performance arrival on coding benchmarks is shaking up the hierarchy of top-tier models.
� Twitter� Reddit
NVIDIA AVO agent achieves 100% on ARC-AGI-3
NVIDIA's AVO agent solved all 183 levels of ARC-AGI-3 without a single pre-programmed goal. It uses a neuro-symbolic approach that separates visual reasoning from standard language prediction, letting it crack unseen, grid-based logic puzzles cold.
Why it matters: A perfect score on ARC-AGI is a massive milestone for truly autonomous, goal-directed AI.
� Twitter� Reddit
🚀 Products & Launches
Grok 4.6 takes #1 spot on CursorBench 3.2
Grok 4.6 topped the CursorBench 3.2 leaderboard with a 70.8% score, proving it's a legitimate powerhouse for IDE-integrated coding. At just $2.81/task, the pricing isn't even close to the competition.
Why it matters: High performance at an aggressive price is a direct threat to every expensive coding model in the room.
� Twitter
Claude Mythos 5 security scans launch in enterprise beta
Anthropic is making a move into the "AI defender" space with Claude Mythos 5, which scans for threats like prompt injection in real time. It includes a transparency dashboard so teams can see exactly why the model flags something.
Why it matters: This is a clear signal that Anthropic is going after enterprise security workflows.
� Twitter
💼 Industry & Business
OpenAI reports 20M active Codex users, offers banked resets
OpenAI confirmed Codex hit 20 million active users and is introducing a "banked reset" policy—unused API quota now rolls over across billing cycles instead of disappearing into the void.
Why it matters: The adoption numbers are massive, and the quota flexibility is a real quality-of-life win for developers building high-demand apps.
� Twitter
GPT-5.6 Sol API prices reduced by over 20%
OpenAI is cutting GPT-5.6 Sol API prices by over 20% for the next three months, covering both input and output tokens. The savings come from improvements in model distillation.
Why it matters: The race to the bottom on price is heating up as everyone fights for dominance in the agentic application layer.
� Twitter
Starcloud secures $2.3B valuation for orbital datacenter mission
Starcloud just hit a $2.3B valuation with a plan to launch satellite clusters that deliver high-performance compute in low-earth orbit for latency-sensitive AI workloads.
Why it matters: The sheer amount of capital behind this shows how far companies are willing to go to solve the AI infrastructure crunch.
� Twitter
Funding & Deals
* Starcloud: Secured a $2.3B valuation to develop orbital datacenters for high-performance AI compute.
* Anthropic: Announced a $35M fund for open-source security research alongside the Claude Mythos 5 release.
Launches
* Grok 4.6: Now available via Vertex AI, sitting at the top of the CursorBench leaderboard with superior coding efficiency.
* Claude Mythos 5: Enterprise beta model featuring real-time security scanning and prompt injection defense.
Closing thought: Between orbital datacenters and agents solving logic puzzles with 100% accuracy, the "frontier" of AI is shifting from just being smart to being physically and architecturally transformative.