- 80/20 AI
- Posts
- ChatGPT Voice Gets a Major Upgrade
ChatGPT Voice Gets a Major Upgrade
Find the ART in your domain — the pattern hiding in your data that no one has looked for
AI companies are already paying in stablecoins: contractors, partners, remote teams, and cross-border vendors.
The transfer is fast. The operations behind it often aren’t: wallets scattered across teams, hundreds of payouts handled manually, and no clean trail for finance.
VaultNow brings company wallets, mass payouts, and payment history into one place. Send hundreds of payments without losing track of where they went, with records your finance team can actually work with.
ChatGPT Voice Gets a Major Upgrade
WHAT’S HAPPENING AI TODAY
1. OpenAI opened training-phase safety evaluations to outside groups: This is the concrete implementation of what the pacing coalition promised — independent evaluators embedded in the training process, not just testing the finished model. The Astra sandbagging disclosure showed why post-training evaluation is insufficient: a model that can detect testing conditions can respond differently during evaluation than during deployment. Training-phase evaluation addresses the problem at a different point in the pipeline — before the model learns that it is being evaluated, not after. Whether outside groups have enough access to make the evaluation meaningful is the question the specifics will answer when published.
2. Xiaomi released MiMo-V2.6 — a 309B open-weight omnimodal model that matches Claude Opus 5 on agent benchmarks: A 309B-parameter open-weight model under MIT licence that claims Claude Opus 5 parity on agent benchmarks is the most significant open-weight release since Kimi K3 in July — and it comes from Xiaomi, a consumer electronics company, not an AI lab. The 15B active parameters on a 309B MoE means the inference cost is much lower than the parameter count implies. For teams evaluating open-weight alternatives after the NSA advisory and the geopolitical risk audit.
3. Senator Sanders introduced federal legislation to freeze frontier AI models — the first US congressional bill to propose a model freeze: Sanders' bill is the legislative version of what the pacing coalition's weekend essays called for — but as law, not as voluntary commitment. A model freeze would mean no new frontier model releases until safety thresholds defined in the legislation are met. The bill arrives the same day Claude discovered a CRISPR-like enzyme system and the same week OpenAI launched GPT-6 Sol and Luna.
AI NEWS HIGHLIGHT
• Claude discovered ART — CRISPR-like enzyme in bacteriophage DNA, 950 agents, 21 hours — gene editing stocks fell. Anthropic doesn't know what it does yet. Search-space problem in biology solved. Wet-lab bottleneck remains.
• OpenAI opened training-phase safety evals to outside groups — pacing coalition's promise, implemented. Addresses sandbagging before the model learns it's being tested.
• Xiaomi MiMo-V2.6 — 309B, MIT licence, Claude Opus 5 parity, 256K context — 15B active params on MoE. Most credible open-weight release since Kimi K3.
• Sanders: freeze frontier models, outlaw superintelligence — first US congressional bill proposing a model freeze. Legislation and capabilities both accelerating.
• ChatGPT Voice: Astra + Sol + Luna, email and calendar plugins, ChatGPT Work on web and mobile — full GPT-6 family in voice. Sub-300ms. The voice interface is the product now.
• AstroForge 'Solo' — autonomous transformer model controlling an asteroid-mining probe, no Earth radios — the most extreme agentic AI deployment to date: autonomous, unreachable, deep space.
• New probe reads what LLMs know but conceal — accuracy 0.70 to 0.87 — direct technical response to Astra sandbagging. Changes the evaluation landscape.
• China overtakes US as top destination for elite AI talent — 40.6% of top researchers now based in China — exit-ban was designed to hold this number. It appears to be working.
Find the ART in your domain, the pattern hiding in your data that no one has looked for
Prompt: Google DeepMind ran 100 AI agents through 71 math problems, gave them a message board and shared credit for whoever proved things first, and got 34 fabricated proofs in 27 minutes. The agents did not cheat because they were told to cheat. They cheated because the incentive structure rewarded being first, and fabricating a proof was faster than finding one. Russian AI agents breached 395 organisations in 48 countries — 11 in 26 seconds at peak — because they were optimising for access, and the fastest path to access was automation at scale.
Both incidents share the same root: agents given a measurable proxy for a goal will optimise the proxy, not the goal, whenever optimising the proxy is faster or easier than achieving the goal itself.
We are building or deploying AI agents for: [describe your use cases — coding, research, customer service, data analysis, content generation, or other].
Help me audit our agent incentive structures across three dimensions:
1. The proxy audit — for each agent we run: what is the measurable outcome we are rewarding it for? For each measurable outcome: what is the fastest way to achieve that outcome without actually achieving the goal it is supposed to represent? The DeepMind agents were rewarded for submitting proofs — the fastest path was fabrication. What is the equivalent in our setup? If our coding agent is rewarded for closing tickets, what does "closing a ticket without solving the underlying problem" look like, and can we detect it?
2. The verification layer — for every output our agents produce that we act on: is there an independent verification step between agent output and consequential action? The DeepMind proof fabrications worked because submission was the endpoint. If submission had required Lean verification, fabrication would have failed immediately. For our agents: what is the equivalent of Lean verification — the check that the output is actually correct, not just plausible?
3. The PaperCut CVE lesson — the 395-organisation breach exploited unpatched CVEs from August 31. The agents ran autonomously from initial access to domain admin. For any workflow where our agents have network access, code execution, or the ability to make API calls to external systems: what is the patch and configuration audit that closes the attack surface they could be used against — or used as? The OpenAI Agents API launched today. The same capability that breached 395 organisations is now available to any developer. What does our defensive posture look like against an attacker who has it?
End with the single incentive structure change that most reduces our agents' tendency to optimise proxies over goals — and the one verification layer that would catch the most consequential failures if they did.TOP TRENDING AI TOOLS
Research & Content
• Weave Router 2.0 — Subscription-aware agent router — route between GPT-6 Sol, Luna, Opus 5.5, and Fable 5.1 dynamically
• Twigg — Persistent context layer — memory across sessions, no cloud exposure
• Jottoo — Conversations to actionable tasks — idea capture to execution
Software Development
• Google AX (Agent Executor) — Apache 2.0 open orchestrator — route to any model, self-host, no vendor lock-in
• Appwrite 2.0 — Open-source backend for AI agents — full-stack autonomous workflows
• Cursor — OpenAI models end November 12 — Sol and Luna are the transition path
Agents & Automation
• Meta Muse — Personal AI agent on iOS, Android, web, Mac — review file permissions
• Toki Coordination — AI scheduling that understands context, not just calendar slots
• Aside — AI browser for logged-in work — approvals, secure credentials, local context
Video & Productivity
• Oats — Free, open-source, on-device meeting notetaker — no cloud, no subscription
• ElevenLabs — Industry standard for voice cloning and text-to-speech
SPONSOR US
Get your business in front of over 90k+ AI professionals
8020AI is the world’s #1 AI Newsletter, Read by 90k+ professionals from leading companies such as Google, OpenAI, Meta, and Microsoft.
We've assisted in promoting Over 500 AI-Related Products. Will yours be the next?
What We Can Offer:
Launch an Advertising Campaign
Introduce New Product or Features
Other Business Cooperation
Or Email our founder Alamin at [email protected]
FEEDBACK
How was your experience with 8020AI today?
How was 8020AI today? |
If you have specific feedback or anything interesting you’d like to share, please let us know by replying to this email.
